Daily digest, 25 Sept 2026
2 top-10 rank changes, 70 new or updated benchmark results across 22 sources.
BenchLeader refreshed 1552 models from 22 of 22 sources. 2 top-10 rank changes, 70 new or updated benchmark results across 22 sources.
Movement in the top 10
- Claude Fable 5.1 (thinking) moved from #5 to #4 in the BenchLeader Index. (details)
- Claude Fable 5.1 (high) moved from #4 to #5 in the BenchLeader Index. (details)
New benchmark results
- ITBench SRE (AA): 10 new results, including GPT-6 Astra, claude-fable-5-1@thinking, claude-opus-5-5@thinking… (board)
- ARC-AGI-1: 3 new results, including Gemini 3.8 Flash, Gemini 3.8 Flash, Gemini 3.8 Flash (board)
- ARC-AGI-2: 3 new results, including Gemini 3.8 Flash, Gemini 3.8 Flash, Gemini 3.8 Flash (board)
- ARC-AGI-3: 3 new results, including Gemini 3.8 Flash, Gemini 3.8 Flash, Gemini 3.8 Flash (board)
- LMArena WebDev: 2 new results, including MiMo-V2.6-Pro, GPT-6 Luna (board)
- AutomationBench: 1 new result, including Claude Opus 5.5 (board)
- GDP.pdf: 1 new result, including Claude Opus 5.5 (board)
- Harvey LAB: 1 new result, including Claude Opus 5.5 (board)
- AA-Briefcase v1.1: 1 new result, including Claude Opus 5.5 (board)
- LMArena Agent: 1 new result, including Grok 4.7 (board)
- EBR-bench: 1 new result, including Claude Opus 5.5 (board)
- FACTS Search: 1 new result, including Claude Opus 5.5 (board)
Score revisions
- Claude Opus 5 (max) on LMArena Agent: 10.2 → 9.5. (details)
- Muse Spark 1.3 (max) on AA-Briefcase: 1597 → 1587. (details)
- GPT-5.6 Sol (xhigh) on LMArena Agent: 7.1 → 6.2. (details)
- Muse Spark 1.3 (xhigh) on AA-Briefcase: 1494 → 1487. (details)
- GPT-5.5 (xhigh) on LMArena Agent: 5 → 4.4. (details)
- Kimi K3 (max) on LMArena Agent: 6.2 → 4.6. (details)
- Qwen3.8 Max (0902) (max) on AA-Briefcase: 1640 → 1626. (details)
- Qwen3.8 Max (0902) (max) on LMArena Agent: 3.3 → 2.8. (details)
- Muse Spark 1.1 on LMArena Agent: -3 → -3.9. (details)
- Qwen3.8 2.4T A95B on AA-Briefcase: 1445 → 1440. (details)
- Gemini 3.8 Flash (high) on LMArena Agent: 4.7 → 4. (details)
- Gemini 3.7 Flash (high) on LMArena Agent: -0.6 → -1.2. (details)
- Grok 4.6 (high) on AA-Briefcase: 1546 → 1539. (details)
- Grok 4.6 (xhigh) on AA-Briefcase: 1555 → 1549. (details)
- Grok 4.6 (xhigh) on LMArena Agent: 2 → 1.2. (details)
- Muse Spark 1.2 (xhigh) on LMArena Agent: -1.8 → -3.1. (details)
- GLM 5.3 Flash on AA-Briefcase: 1459 → 1452. (details)
- GLM 5.3 Flash on LMArena Agent: 1.1 → 0.3. (details)
- Gemini 3.1 Pro on LMArena Agent: -5.8 → -6.8. (details)
- GPT-5.5 on LMArena Agent: 2.7 → 1.8. (details)
- …and 22 more.
Speed changes
- GPT-5.6 Sol time to first token changed from 2.52 s to 3.63 s. (details)
- GPT-6 Sol time to first answer changed from 8.89 s to 12.22 s. (details)
- GLM 5.3 time to first token changed from 0.68 s to 1.16 s. (details)
- Muse Spark 1.1 time to first token changed from 1.64 s to 2.70 s. (details)
- Qwen3.8 2.4T A95B time to first token changed from 2.00 s to 1.04 s. (details)
- GPT-5.4 Pro output speed changed from 1 tok/s to 2 tok/s. (details)
- Kimi K2.5 time to first token changed from 2.33 s to 1.48 s. (details)
- GLM 5.1 output speed changed from 37 tok/s to 58 tok/s. (details)
- GLM 5.1 thinking time changed from 101.31 s to 65.53 s. (details)
- GPT-5.1-Codex time to first token changed from 10.80 s to 6.49 s. (details)
- GLM 5V Turbo time to first token changed from 3.08 s to 4.66 s. (details)
- DeepSeek V3.2 output speed changed from 3 tok/s to 5 tok/s. (details)
- Muse Glimmer time to first token changed from 0.52 s to 0.75 s. (details)
- Qwen3.5-27B time to first token changed from 2.44 s to 0.87 s. (details)
- Solar Pro 4 output speed changed from 59 tok/s to 81 tok/s. (details)
- …and 55 more.
Today's top five
- Claude Opus 5.5 — 72.0
- Claude Fable 5.1 — 70.8
- GPT-6 Astra — 70.2
- Claude Opus 5.5 — 70.2
- Gemini 4 Argon — 70.0
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.