BenchLeader

BALROG

Progress through classic text and grid games such as NetHack.

As of 19 Sept 2026, Gemini 3 Pro leads BALROG on BenchLeader with 58.1%, ahead of Gemini 3.1 Pro at 57.0%, across 32 model configurations with a published result.

Published by
BALROGdata via Epoch AI Benchmarking Hub
Category
Agents & tools
Index weight
Reference only
Models
32
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Agents play games including BabyAI, Crafter, TextWorld and NetHack; the score is average progress toward each game's goal.

How it is scored

Average progress across games, as published by BALROG.

What to keep in mind

Progress is normalised per game and games differ enormously in difficulty.

32 of 32
#
1Gemini 3 ProGoogle58.1%61.12025-11-18
2Gemini 3.1 ProGoogle57.0%63.92026-02-19
3Gemini 3 FlashGoogle48.1%57.72025-12-17
4Grok 4xAI43.6%58.52025-07-09
5Claude Opus 4.5Anthropic43.5%58.22025-11-24
6Gemini 2.5 ProGoogle43.3%54.52025-03-25
7DeepSeek R1DeepSeekopen34.9%48.72025-01-20
8Gemini 2.5 FlashGoogle33.5%52.52025-06-17
9GPT-5minimalOpenAI32.8%46.02025-08-07
10Claude 3.5 SonnetAnthropic32.6%46.62024-10-22
11GPT-4oOpenAI32.3%44.32024-05-13
12Claude Haiku 4.5Anthropic31.2%48.12025-10-15
13Grok 3xAI29.5%51.02025-04-09
14Reka Flash 3Rekaopen ↗29.2%39.32025-03-10
15Llama 3.1 70BMetaopen27.9%41.32024-07-23
16Llama 3.2 90BMetaopen ↗27.3%37.62024-09-24
17Llama 3.3 70BMetaopen23.0%41.42024-12-06
18Gemini 1.5 Pro 002Google21.0%44.92024-09-24
19DeepSeek-R1-Distill-Qwen-32BDeepSeekopen ↗19.5%42.92025-01-20
20Claude 3.5 HaikuAnthropic19.3%40.32024-10-22
21Mistral NemoMistral AIopen ↗17.6%2024-07-18
22GPT-4o miniOpenAI17.4%36.22024-07-18
23Llama 3.2 Instruct 11B (Vision)Metaopen ↗16.8%37.12024-09-24
24Qwen2.5 72BAlibabaopen16.2%42.72024-09-19
25Llama 3.1 8BMetaopen15.1%36.82024-07-23
26Gemini 1.5 Flash 002Google14.6%40.82024-09-24
27Qwen2 Vl 72BAlibaba12.8%2024-08-29
28Phi-4Microsoftopen11.6%39.62024-12-12
29Llama 3.2 3B InstructMetaopen ↗10.1%36.32024-09-24
30Qwen2.5 7B InstructAlibaba7.8%40.12024-09-19
31Llama 3.2 1B InstructMetaopen ↗6.6%33.02024-09-24
32Qwen2 Vl 7BAlibaba3.7%2024-08-29

Cite as: BenchLeader, “BALROG leaderboard”, https://www.benchleader.com/benchmarks/balrog, data as of 19 Sept 2026.

BALROG: questions

What does BALROG measure?
Agents play games including BabyAI, Crafter, TextWorld and NetHack; the score is average progress toward each game's goal. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads BALROG?
Gemini 3 Pro leads BALROG with 58.1% as of 19 Sept 2026, ahead of Gemini 3.1 Pro at 57.0%.
How many models have BALROG results?
32 model configurations have a BALROG result on BenchLeader, all taken from BALROG via Epoch AI Benchmarking Hub.
Who runs BALROG and how often is it updated?
BALROG is published by BALROG. BenchLeader re-reads the published results every morning and records the date each result was published.
Does BALROG count toward the BenchLeader Index?
No. BALROG is shown for reference but left out of the composite index.