Daily digest, 9 Oct 2026
2 new models, 3 top-10 rank changes, 3 price changes, 151 new or updated benchmark results across 23 sources.
BenchLeader refreshed 1561 models from 23 of 23 sources. 2 new models, 3 top-10 rank changes, 3 price changes, 151 new or updated benchmark results across 23 sources.
New models
- GPT-6 Sol (Daybreak Blue, max) (max) (OpenAI) is now listed at $4.00/M blended; no independent results yet. (details)
- Ling 3.0 Flash Sante (Ant Group) is now listed at $0.06/M blended; no independent results yet. (details)
Movement in the top 10
- Gemini 4 Argon (high) moved from #7 to #5 in the BenchLeader Index. (details)
- GPT-6 Astra (high) moved from #5 to #6 in the BenchLeader Index. (details)
- Claude Fable 5.1 (xhigh) moved from #6 to #7 in the BenchLeader Index. (details)
Price changes
- Step 5 Preview cost per task moved from 0.7 to 1 (43% up). (details)
- Step 5 Preview (high) cost per task moved from 0.7 to 1 (43% up). (details)
- Nano Banana 2.1 output price moved from $7.50/M to $30.00/M (300% up). (details)
New benchmark results
- AudioMC (text output): 17 new results, including Inkling Small, Gemini 2.5 Flash, Gemini 2.5 Flash… (board)
- Harvey LAB: 16 new results, including GPT-6 Astra, GPT-6.1 Sol, Muse Spark 1.3… (board)
- SciPredict: 15 new results, including Gemini 3 Pro, o3, GPT-5.2… (board)
- AudioMC (audio output): 15 new results, including Qwen3 Omni 30B A3B Instruct, Gemini 2.5 Flash Native Audio 12 2025, Gemini 2.5 Flash Native Audio 12 2025… (board)
- AA-Briefcase v1.1: 6 new results, including Gemini 3.1 Pro, Claude Haiku 5.5, Claude Haiku 5.5… (board)
- AutomationBench: 5 new results, including Gemini 3.1 Pro, Claude Haiku 5.5, Claude Haiku 5.5… (board)
- GDP.pdf: 5 new results, including Gemini 3.1 Pro, Claude Haiku 5.5, Claude Haiku 5.5… (board)
- LMArena Maths: 4 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview… (board)
- LMArena Agent: 2 new results, including MiMo-V2.6-Pro, Mistral Large 4 (board)
- MLCR: 2 new results, including Gemini 3.1 Pro, Claude Haiku 5.5 (board)
- LMArena WebDev: 2 new results, including Claude Haiku 5.5, Mistral Large 4 (board)
- EnterpriseOps-Gym: 1 new result, including Gemini 3.1 Pro (board)
Score revisions
- Claude Opus 5.5 (max) on Harvey LAB: 91.2% → 4.2%. (details)
- Claude Fable 5.1 (max) on Harvey LAB: 93% → 6.4%. (details)
- Claude Fable 5.1 (max) on LMArena Agent: 14.3 → 12.7. (details)
- GPT-6 Astra (max) on LMArena Agent: 12.3 → 13.1. (details)
- Claude Opus 5.5 (high) on LMArena Coding: 1539 → 1550. (details)
- Claude Opus 5.5 (high) on LMArena Agent: 13.8 → 14.3. (details)
- Claude Opus 5.5 (high) on LMArena Maths: 1511 → 1499. (details)
- Gemini 4 Argon (high) on LMArena Agent: 7.6 → 9.3. (details)
- GPT-6.1 Sol (max) on LMArena Agent: 11.2 → 11.7. (details)
- GPT-6.1 Sol (max) on LMArena Multi-turn: 1493 → 1487. (details)
- Claude Opus 5 (high) on LMArena Agent: 8.7 → 8. (details)
- Muse Spark 1.3 (max) on LMArena Maths: 1509 → 1502. (details)
- Claude Sonnet 5.5 (max) on Harvey LAB: 93.1% → 2.8%. (details)
- Claude Sonnet 5.5 (max) on LMArena Agent: 12.5 → 12. (details)
- Claude Sonnet 5.5 (xhigh) on LMArena Creative Writing: 1447 → 1463. (details)
- Claude Sonnet 5.5 (xhigh) on LMArena Instruction Following: 1484 → 1489. (details)
- Claude Sonnet 5.5 (xhigh) on LMArena Multi-turn: 1473 → 1484. (details)
- Claude Sonnet 5.5 (xhigh) on LMArena Longer Queries: 1493 → 1503. (details)
- GPT-6 Sol (max) on LMArena Creative Writing: 1442 → 1436. (details)
- Kimi K3 (max) on Harvey LAB: 94.6% → 5.3%. (details)
- …and 32 more.
Speed changes
- GPT-6 Astra time to first token changed from 9.55 s to 4.00 s. (details)
- Claude Opus 5 time to first answer changed from 11.69 s to 17.07 s. (details)
- Claude Opus 5 time to first token changed from 1.42 s to 3.10 s. (details)
- GPT-5.6 Sol time to first answer changed from 134.89 s to 85.27 s. (details)
- Claude Fable 5.1 time to first answer changed from 7.60 s to 4.51 s. (details)
- Claude Opus 5 output speed changed from 46 tok/s to 67 tok/s. (details)
- GPT-5.5 Pro output speed changed from 28 tok/s to 15 tok/s. (details)
- MiMo-V2.6-Pro time to first token changed from 9.47 s to 3.44 s. (details)
- Grok 4.7 time to first answer changed from 87.05 s to 50.24 s. (details)
- Grok 4.7 response time changed from 93.42 s to 57.99 s. (details)
- GLM 5.3 time to first token changed from 1.82 s to 1.08 s. (details)
- Claude Opus 4.7 time to first token changed from 3.13 s to 1.77 s. (details)
- Qwen3.8 2.4T A95B time to first token changed from 4.02 s to 1.48 s. (details)
- GPT-5.5 time to first answer changed from 8.29 s to 5.30 s. (details)
- Muse Spark 1.1 output speed changed from 153 tok/s to 93 tok/s. (details)
- …and 72 more.
Today's top five
- Claude Opus 5.5 — 72.0
- Claude Fable 5.1 — 70.8
- GPT-6 Astra — 70.2
- Claude Opus 5.5 — 70.2
- Gemini 4 Argon — 70.0
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.