BALROG
Progress through classic text and grid games such as NetHack.
As of 19 Sept 2026, Gemini 3 Pro leads BALROG on BenchLeader with 58.1%, ahead of Gemini 3.1 Pro at 57.0%, across 32 model configurations with a published result.
- Published by
- BALROGdata via Epoch AI Benchmarking Hub
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 32
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Agents play games including BabyAI, Crafter, TextWorld and NetHack; the score is average progress toward each game's goal.
How it is scored
Average progress across games, as published by BALROG.
What to keep in mind
Progress is normalised per game and games differ enormously in difficulty.
- 1Gemini 3 Pro58.1%
- 2Gemini 3.1 Pro57.0%
- 3Gemini 3 Flash48.1%
- 4Grok 443.6%
- 5Claude Opus 4.543.5%
- 6Gemini 2.5 Pro43.3%
- 7DeepSeek R134.9%
- 8Gemini 2.5 Flash33.5%
- 9GPT-5 (minimal)32.8%
- 10Claude 3.5 Sonnet32.6%
- 11GPT-4o32.3%
- 12Claude Haiku 4.531.2%
- 13Grok 329.5%
- 14Reka Flash 329.2%
- 15Llama 3.1 70B27.9%
32 of 32
| # | ||||
|---|---|---|---|---|
| 1 | 58.1% | 61.1 | 2025-11-18 | |
| 2 | 57.0% | 63.9 | 2026-02-19 | |
| 3 | 48.1% | 57.7 | 2025-12-17 | |
| 4 | 43.6% | 58.5 | 2025-07-09 | |
| 5 | 43.5% | 58.2 | 2025-11-24 | |
| 6 | 43.3% | 54.5 | 2025-03-25 | |
| 7 | 34.9% | 48.7 | 2025-01-20 | |
| 8 | 33.5% | 52.5 | 2025-06-17 | |
| 9 | 32.8% | 46.0 | 2025-08-07 | |
| 10 | 32.6% | 46.6 | 2024-10-22 | |
| 11 | 32.3% | 44.3 | 2024-05-13 | |
| 12 | 31.2% | 48.1 | 2025-10-15 | |
| 13 | 29.5% | 51.0 | 2025-04-09 | |
| 14 | 29.2% | 39.3 | 2025-03-10 | |
| 15 | 27.9% | 41.3 | 2024-07-23 | |
| 16 | 27.3% | 37.6 | 2024-09-24 | |
| 17 | 23.0% | 41.4 | 2024-12-06 | |
| 18 | 21.0% | 44.9 | 2024-09-24 | |
| 19 | 19.5% | 42.9 | 2025-01-20 | |
| 20 | 19.3% | 40.3 | 2024-10-22 | |
| 21 | 17.6% | – | 2024-07-18 | |
| 22 | 17.4% | 36.2 | 2024-07-18 | |
| 23 | 16.8% | 37.1 | 2024-09-24 | |
| 24 | 16.2% | 42.7 | 2024-09-19 | |
| 25 | 15.1% | 36.8 | 2024-07-23 | |
| 26 | 14.6% | 40.8 | 2024-09-24 | |
| 27 | 12.8% | – | 2024-08-29 | |
| 28 | 11.6% | 39.6 | 2024-12-12 | |
| 29 | 10.1% | 36.3 | 2024-09-24 | |
| 30 | 7.8% | 40.1 | 2024-09-19 | |
| 31 | 6.6% | 33.0 | 2024-09-24 | |
| 32 | 3.7% | – | 2024-08-29 |
Cite as: BenchLeader, “BALROG leaderboard”, https://www.benchleader.com/benchmarks/balrog, data as of 19 Sept 2026.
BALROG: questions
- What does BALROG measure?
- Agents play games including BabyAI, Crafter, TextWorld and NetHack; the score is average progress toward each game's goal. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads BALROG?
- Gemini 3 Pro leads BALROG with 58.1% as of 19 Sept 2026, ahead of Gemini 3.1 Pro at 57.0%.
- How many models have BALROG results?
- 32 model configurations have a BALROG result on BenchLeader, all taken from BALROG via Epoch AI Benchmarking Hub.
- Who runs BALROG and how often is it updated?
- BALROG is published by BALROG. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does BALROG count toward the BenchLeader Index?
- No. BALROG is shown for reference but left out of the composite index.