BenchLeader

GBAEval

Playing Game Boy Advance games as an agent.

As of 19 Sept 2026, Claude Opus 5 leads GBAEval on BenchLeader with 79.6%, ahead of Claude Fable 5 at 74.5%, across 23 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Agents & tools
Index weight
Reference only
Models
23
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

The agent plays GBA games from screen and controls, scored on progress.

How it is scored

Overall score, as published.

What to keep in mind

A games benchmark; small set.

23 of 23
#
1Claude Opus 5Anthropic79.6%66.72026-07-24
2Claude Fable 5Anthropic74.5%68.32026-06-09
3Claude Opus 4.8Anthropic70.9%61.92026-05-28
4Grok 4.5xAI65.4%60.32026-07-08
5Claude Sonnet 5Anthropic65.3%57.22026-06-30
6GPT-5.5OpenAI53.2%63.22026-04-23
7GPT-5.6 SolOpenAI52.6%56.82026-07-09
8Claude Sonnet 4.6Anthropic48.8%59.02026-02-17
9Kimi K3Moonshot AIopen ↗48.4%64.22026-07-16
10GPT-5.4OpenAI45.1%59.22026-03-05
11Claude Opus 4.6Anthropic44.1%63.72026-02-05
12Claude Opus 4.7Anthropic43.8%64.52026-04-16
13Muse Spark 1.1Meta7.9%65.22026-07-09
14Gemini 3.5 FlashGoogle6.7%51.32026-05-19
15Grok Build 0.1xAI2.4%54.92026-05-29
16MiniMax-M3MiniMaxopen ↗0.9%57.52026-06-01
17Kimi K2.6Moonshot AIopen ↗0.9%60.52026-04-20
18Gemini 3.1 ProGoogle0.8%63.92026-02-19
19Kimi K2.7 CodeMoonshot AIopen ↗0.8%55.92026-06-12
20Qwen3 7maxAlibaba0.4%61.02026-05-19
21GLM-5.2Zhipu AIopen ↗0.0%52.12026-06-16
22GLM-5.1Zhipu AIopen ↗0.0%56.92026-04-07
23MiniMax-M2.7MiniMaxopen ↗0.0%56.12026-03-18

Cite as: BenchLeader, “GBAEval leaderboard”, https://www.benchleader.com/benchmarks/gbaeval, data as of 19 Sept 2026.

GBAEval: questions

What does GBAEval measure?
The agent plays GBA games from screen and controls, scored on progress. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads GBAEval?
Claude Opus 5 leads GBAEval with 79.6% as of 19 Sept 2026, ahead of Claude Fable 5 at 74.5%.
How many models have GBAEval results?
23 model configurations have a GBAEval result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs GBAEval and how often is it updated?
GBAEval is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does GBAEval count toward the BenchLeader Index?
No. GBAEval is shown for reference but left out of the composite index.