GBAEval
Playing Game Boy Advance games as an agent.
As of 19 Sept 2026, Claude Opus 5 leads GBAEval on BenchLeader with 79.6%, ahead of Claude Fable 5 at 74.5%, across 23 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 23
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
The agent plays GBA games from screen and controls, scored on progress.
How it is scored
Overall score, as published.
What to keep in mind
A games benchmark; small set.
- 1Claude Opus 579.6%
- 2Claude Fable 574.5%
- 3Claude Opus 4.870.9%
- 4Grok 4.565.4%
- 5Claude Sonnet 565.3%
- 6GPT-5.553.2%
- 7GPT-5.6 Sol52.6%
- 8Claude Sonnet 4.648.8%
- 9Kimi K348.4%
- 10GPT-5.445.1%
- 11Claude Opus 4.644.1%
- 12Claude Opus 4.743.8%
- 13Muse Spark 1.17.9%
- 14Gemini 3.5 Flash6.7%
- 15Grok Build 0.12.4%
23 of 23
| # | ||||
|---|---|---|---|---|
| 1 | 79.6% | 66.7 | 2026-07-24 | |
| 2 | 74.5% | 68.3 | 2026-06-09 | |
| 3 | 70.9% | 61.9 | 2026-05-28 | |
| 4 | 65.4% | 60.3 | 2026-07-08 | |
| 5 | 65.3% | 57.2 | 2026-06-30 | |
| 6 | 53.2% | 63.2 | 2026-04-23 | |
| 7 | 52.6% | 56.8 | 2026-07-09 | |
| 8 | 48.8% | 59.0 | 2026-02-17 | |
| 9 | 48.4% | 64.2 | 2026-07-16 | |
| 10 | 45.1% | 59.2 | 2026-03-05 | |
| 11 | 44.1% | 63.7 | 2026-02-05 | |
| 12 | 43.8% | 64.5 | 2026-04-16 | |
| 13 | 7.9% | 65.2 | 2026-07-09 | |
| 14 | 6.7% | 51.3 | 2026-05-19 | |
| 15 | 2.4% | 54.9 | 2026-05-29 | |
| 16 | 0.9% | 57.5 | 2026-06-01 | |
| 17 | 0.9% | 60.5 | 2026-04-20 | |
| 18 | 0.8% | 63.9 | 2026-02-19 | |
| 19 | 0.8% | 55.9 | 2026-06-12 | |
| 20 | 0.4% | 61.0 | 2026-05-19 | |
| 21 | 0.0% | 52.1 | 2026-06-16 | |
| 22 | 0.0% | 56.9 | 2026-04-07 | |
| 23 | 0.0% | 56.1 | 2026-03-18 |
Cite as: BenchLeader, “GBAEval leaderboard”, https://www.benchleader.com/benchmarks/gbaeval, data as of 19 Sept 2026.
GBAEval: questions
- What does GBAEval measure?
- The agent plays GBA games from screen and controls, scored on progress. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads GBAEval?
- Claude Opus 5 leads GBAEval with 79.6% as of 19 Sept 2026, ahead of Claude Fable 5 at 74.5%.
- How many models have GBAEval results?
- 23 model configurations have a GBAEval result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs GBAEval and how often is it updated?
- GBAEval is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does GBAEval count toward the BenchLeader Index?
- No. GBAEval is shown for reference but left out of the composite index.