Daily digest, 3 Oct 2026
1 new model, 9 top-10 rank changes, 22 price changes, 149 new or updated benchmark results across 23 sources.
BenchLeader refreshed 1550 models from 23 of 23 sources. 1 new model, 9 top-10 rank changes, 22 price changes, 149 new or updated benchmark results across 23 sources.
New models
- Ling 3.1 Flash (Ant Group) is now listed; no independent results yet. (details)
Movement in the top 10
- Claude Opus 5.5 (max) moved from #4 to #3 in the BenchLeader Index. (details)
- Claude Fable 5.1 (xhigh) moved from #5 to #4 in the BenchLeader Index. (details)
- GPT-6 Astra (high) moved from #3 to #5 in the BenchLeader Index. (details)
- Claude Fable 5.1 (high) moved from #7 to #6 in the BenchLeader Index. (details)
- Claude Opus 5.5 (high) moved from #6 to #7 in the BenchLeader Index. (details)
- Claude Opus 5.5 (xhigh) moved from #9 to #8 in the BenchLeader Index. (details)
- Claude Fable 5.1 (thinking) moved from #8 to #9 in the BenchLeader Index. (details)
- GPT-6.1 Sol (max) moved from #11 to #10 in the BenchLeader Index. (details)
- Gemini 4 Argon (high) moved from #10 to #11 in the BenchLeader Index. (details)
Price changes
- DeepSeek V4 Pro (max) input price moved from $0.43/M to $0.66/M (52% up). (details)
- DeepSeek V4 Pro (max) output price moved from $0.87/M to $1.98/M (128% up). (details)
- DeepSeek V4 Pro (max) cached input price moved from $0.00/M to $0.02/M (507% up). (details)
- DeepSeek V4 Pro (high) input price moved from $0.43/M to $0.66/M (52% up). (details)
- DeepSeek V4 Pro (high) output price moved from $0.87/M to $1.98/M (128% up). (details)
- DeepSeek V4 Pro (high) cached input price moved from $0.00/M to $0.02/M (507% up). (details)
- MiMo-V2.5-Pro input price moved from $0.52/M to $1.00/M (92% up). (details)
- MiMo-V2.5-Pro output price moved from $1.04/M to $3.00/M (187% up). (details)
- DeepSeek V4 Pro input price moved from $0.43/M to $0.66/M (52% up). (details)
- DeepSeek V4 Pro output price moved from $0.87/M to $1.98/M (128% up). (details)
- DeepSeek V4 Pro cached input price moved from $0.00/M to $0.02/M (507% up). (details)
- MiMo-V2.5-Pro (no reasoning) input price moved from $0.52/M to $1.00/M (92% up). (details)
- MiMo-V2.5-Pro (no reasoning) output price moved from $1.04/M to $3.00/M (187% up). (details)
- DeepSeek V4 Pro (no reasoning) input price moved from $0.43/M to $0.66/M (52% up). (details)
- DeepSeek V4 Pro (no reasoning) output price moved from $0.87/M to $1.98/M (128% up). (details)
- DeepSeek V4 Pro (no reasoning) cached input price moved from $0.00/M to $0.02/M (507% up). (details)
- DeepSeek V4 Pro (low) input price moved from $0.43/M to $0.66/M (52% up). (details)
- DeepSeek V4 Pro (low) output price moved from $0.87/M to $1.98/M (128% up). (details)
- DeepSeek V4 Pro (low) cached input price moved from $0.00/M to $0.02/M (507% up). (details)
- DeepSeek V4 Pro (xhigh) input price moved from $0.43/M to $0.66/M (52% up). (details)
- DeepSeek V4 Pro (xhigh) output price moved from $0.87/M to $1.98/M (128% up). (details)
- DeepSeek V4 Pro (xhigh) cached input price moved from $0.00/M to $0.02/M (507% up). (details)
New benchmark results
- MultiChallenge: 28 new results, including Muse Spark, Muse Spark 1.1, GPT-5.4 Pro… (board)
- LMArena Agent: 4 new results, including GPT-6.1 Sol, Step 5 Preview, Claude Sonnet 5.5… (board)
- LMArena Text: 3 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview (board)
- LMArena Hard Prompts: 3 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview (board)
- LMArena Coding: 3 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview (board)
- LMArena Vision: 3 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview (board)
- LMArena Creative Writing: 3 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview (board)
- LMArena Instruction Following: 3 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview (board)
- LMArena Multi-turn: 3 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview (board)
- LMArena Longer Queries: 3 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview (board)
- AA Intelligence Index v4.3.2: 3 new results, including Grok 4.7, Claude Sonnet 5.5, Kimi K2.6 (board)
- SciCode (AA): 3 new results, including Grok 4.7, Kimi K2.6, Gemma 4 E4B (board)
Score revisions
- Claude Opus 5.5 (thinking) on AA-Briefcase: 1822 → 1808. (details)
- Claude Fable 5.1 (xhigh) on AA-Briefcase: 1669 → 1657. (details)
- Claude Fable 5.1 (high) on AA-Briefcase: 1592 → 1581. (details)
- Claude Opus 5.5 (high) on AA-Briefcase: 1705 → 1690. (details)
- Claude Opus 5.5 (xhigh) on AA-Briefcase: 1780 → 1768. (details)
- Claude Opus 5 (high) on AA-Briefcase: 1573 → 1562. (details)
- Claude Opus 5 (max) on AA-Briefcase: 1673 → 1662. (details)
- Claude Opus 5 (max) on LMArena Agent: 8.5 → 7.9. (details)
- Claude Opus 5 (xhigh) on AA-Briefcase: 1649 → 1636. (details)
- GPT-5.6 Sol (xhigh) on LMArena Agent: 7.1 → 6.5. (details)
- GPT-6 Sol (max) on GDPval (AA): 49.4% → 50.3%. (details)
- GPT-6 Sol (max) on LMArena Agent: 10.6 → 9.7. (details)
- Claude Sonnet 5.5 (high) on GDPval (AA): 50.9% → 52.6%. (details)
- GPT-5.5 (high) on GDPval (AA): 40.9% → 41.5%. (details)
- Claude Opus 5 (medium) on AA-Briefcase: 1441 → 1435. (details)
- GPT-6 Sol (xhigh) on GDPval (AA): 46.8% → 47.8%. (details)
- GLM 5.3 (max) on GDPval (AA): 57.2% → 57.8%. (details)
- GPT-6 Sol (high) on MMMU-Pro: 81.2% → 82%. (details)
- GPT-6 Sol (high) on GDPval (AA): 43.8% → 44.8%. (details)
- GPT-5.6 Terra (max) on GDPval (AA): 46.6% → 47.7%. (details)
- …and 37 more.
Speed changes
- GPT-5.5 time to first answer changed from 15.32 s to 21.23 s. (details)
- GPT-6 Sol output speed changed from 63 tok/s to 87 tok/s. (details)
- Grok 4.7 time to first answer changed from 78.69 s to 49.32 s. (details)
- Grok 4.7 response time changed from 85.56 s to 55.49 s. (details)
- Muse Spark 1.1 output speed changed from 215 tok/s to 124 tok/s. (details)
- Claude Sonnet 5 time to first answer changed from 16.43 s to 25.45 s. (details)
- Kimi K2.6 output speed changed from 69 tok/s to 40 tok/s. (details)
- Ling 3.0 Flash VL output speed changed from 142 tok/s to 48 tok/s. (details)
- DeepSeek V3.2 output speed changed from 29 tok/s to 15 tok/s. (details)
- MiMo-V2.5 time to first answer changed from 58.82 s to 38.10 s. (details)
- Solar Pro 4 output speed changed from 80 tok/s to 109 tok/s. (details)
- Grok Build 0.1 output speed changed from 48 tok/s to 65 tok/s. (details)
- GPT-5 time to first answer changed from 7.89 s to 15.19 s. (details)
- GPT-5 response time changed from 15.03 s to 21.78 s. (details)
- Qwen3.5 397B A17B output speed changed from 39 tok/s to 53 tok/s. (details)
- …and 17 more.
Today's top five
- Claude Opus 5.5 — 72.0
- Claude Fable 5.1 — 70.8
- GPT-6 Astra — 70.2
- Claude Opus 5.5 — 70.2
- Gemini 4 Argon — 70.0
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.