Daily digest, 14 Sept 2026
8 top-10 rank changes, 12 new or updated benchmark results across 20 sources.
BenchLeader refreshed 1355 models from 20 of 20 sources. 8 top-10 rank changes, 12 new or updated benchmark results across 20 sources.
Movement in the top 10
- Claude Fable 5.1 (xhigh) moved from #3 to #2 in the BenchLeader Index. (details)
- Claude Fable 5 moved from #4 to #3 in the BenchLeader Index. (details)
- GPT-6 Astra (high) moved from #5 to #4 in the BenchLeader Index. (details)
- Claude Fable 5.1 (high) moved from #7 to #5 in the BenchLeader Index. (details)
- GPT-6 Astra moved from #8 to #6 in the BenchLeader Index. (details)
- Claude Fable 5.1 (max) moved from #6 to #7 in the BenchLeader Index. (details)
- Claude Opus 5 moved from #9 to #8 in the BenchLeader Index. (details)
- GPT-6 Astra (max) moved from #2 to #9 in the BenchLeader Index. (details)
New benchmark results
- LMArena Vision: 4 new results, including Claude Fable 5.1, GPT-6 Astra, Muse Spark 1.3… (board)
- PRBench Finance: 1 new result, including Muse Spark 1.3 (board)
- PRBench Legal: 1 new result, including Muse Spark 1.3 (board)
- Epoch Capabilities Index: 1 new result, including Gemini 3.8 Flash (board)
- LMArena Text: 1 new result, including Muse Spark 1.3 (board)
- LMArena Hard Prompts: 1 new result, including Muse Spark 1.3 (board)
- LMArena Coding: 1 new result, including Muse Spark 1.3 (board)
Score revisions
- GPT-6 Astra (max) on LMArena Coding: 1537 → 1543. (details)
- Qwen3.8 27B on LMArena Vision: 1279 → 1272. (details)
Speed changes
- Gemini 3.8 Flash time to first answer changed from 12.81 s to 17.67 s. (details)
- Nemotron 3 Ultra 550B A55B time to first token changed from 1.09 s to 0.40 s. (details)
- Ling Flash 2.0 output speed changed from 3 tok/s to 5 tok/s. (details)
- Llama 3.1 Nemotron 70B Instruct output speed changed from 60 tok/s to 91 tok/s. (details)
- Llama 3.1 Nemotron 70B Instruct time to first answer changed from 8.36 s to 4.33 s. (details)
- Llama 3.1 Nemotron 70B Instruct time to first token changed from 8.36 s to 4.33 s. (details)
- Nemotron Nano 12B v2 VL output speed changed from 54 tok/s to 82 tok/s. (details)
- Qwen2.5 Coder 32B Instruct time to first token changed from 0.98 s to 1.40 s. (details)
- Granite 4.0 H Small output speed changed from 20 tok/s to 12 tok/s. (details)
- Nemotron 3.5 Lightning time to first token changed from 0.62 s to 0.40 s. (details)
- Lyria 3 Pro time to first token changed from 3.30 s to 2.12 s. (details)
Today's top five
- Claude Fable 5.1 — 70.6
- Claude Fable 5.1 — 70.4
- Claude Fable 5 — 70.2
- GPT-6 Astra — 70.2
- Claude Fable 5.1 — 70.0
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.