Daily digest, 8 Oct 2026
5 new models, 11 top-10 rank changes, 26 price changes, 157 new or updated benchmark results across 23 sources.
BenchLeader refreshed 1557 models from 23 of 23 sources. 5 new models, 11 top-10 rank changes, 26 price changes, 157 new or updated benchmark results across 23 sources.
New models
- Claude Haiku 5.5 (xhigh) (Anthropic) appeared with 5 benchmark results, entering the BenchLeader Index at #104. (details)
- Claude Haiku 5.5 (max) (Anthropic) appeared with 5 benchmark results, entering the BenchLeader Index at #121. (details)
- Claude Haiku 5.5 (low) (Anthropic) appeared with 4 benchmark results. (details)
- Claude Haiku 5.5 (medium) (Anthropic) appeared with 4 benchmark results. (details)
- Claude Haiku 5.5 (high) (Anthropic) appeared with 4 benchmark results. (details)
Movement in the top 10
- Claude Opus 5.5 (max) moved from #10 to #1 in the BenchLeader Index. (details)
- Claude Fable 5.1 (max) moved from #14 to #2 in the BenchLeader Index. (details)
- Claude Opus 5.5 (high) moved from #5 to #4 in the BenchLeader Index. (details)
- GPT-6 Astra (high) moved from #6 to #5 in the BenchLeader Index. (details)
- Claude Fable 5.1 (xhigh) moved from #7 to #6 in the BenchLeader Index. (details)
- Gemini 4 Argon (high) moved from #8 to #7 in the BenchLeader Index. (details)
- Claude Fable 5.1 (high) moved from #9 to #8 in the BenchLeader Index. (details)
- Claude Opus 5.5 (xhigh) moved from #11 to #9 in the BenchLeader Index. (details)
- GPT-6.1 Sol (max) moved from #12 to #10 in the BenchLeader Index. (details)
- Claude Fable 5 moved from #4 to #18 in the BenchLeader Index. (details)
- Claude Fable 5.1 moved from #1 to #34 in the BenchLeader Index. (details)
Price changes
- Claude Opus 5.5 (high) cost per task moved from 6 to 1.8 (70% down). (details)
- Claude Fable 5.1 (xhigh) cost per task moved from 7.6 to 6 (22% down). (details)
- Claude Fable 5.1 (high) cost per task moved from 7.6 to 3.9 (49% down). (details)
- Claude Opus 5.5 (xhigh) cost per task moved from 6 to 3.5 (42% down). (details)
- Claude Sonnet 5.5 (max) cost per task moved from 7.7 to 5.5 (29% down). (details)
- Claude Sonnet 5.5 (max) cached input price moved from $0.20/M to $0.10/M (50% down). (details)
- Claude Fable 5.1 (medium) cost per task moved from 7.6 to 3 (61% down). (details)
- Claude Sonnet 5.5 (xhigh) cost per task moved from 7.7 to 2 (74% down). (details)
- Claude Sonnet 5.5 (xhigh) cached input price moved from $0.20/M to $0.10/M (50% down). (details)
- Claude Opus 5.5 (medium) cost per task moved from 6 to 1.3 (78% down). (details)
- Claude Sonnet 5.5 (high) cost per task moved from 7.7 to 0.9 (88% down). (details)
- Claude Sonnet 5.5 (high) cached input price moved from $0.20/M to $0.10/M (50% down). (details)
- Claude Fable 5.1 (low) cost per task moved from 7.6 to 2.4 (69% down). (details)
- Claude Opus 5.5 (low) cost per task moved from 6 to 0.6 (91% down). (details)
- Claude Sonnet 5.5 (medium) cost per task moved from 7.7 to 0.5 (94% down). (details)
- Claude Sonnet 5.5 (medium) cached input price moved from $0.20/M to $0.10/M (50% down). (details)
- Claude Sonnet 5.5 (low) cost per task moved from 7.7 to 0.3 (95% down). (details)
- Claude Sonnet 5.5 (low) cached input price moved from $0.20/M to $0.10/M (50% down). (details)
- DeepSeek V4 Flash (high) input price moved from $0.11/M to $0.13/M (18% up). (details)
- DeepSeek V4 Flash (high) output price moved from $0.24/M to $0.28/M (17% up). (details)
- Nemotron 3 Nano 30B A3B input price moved from $0.05/M to $0.06/M (20% up). (details)
- Nemotron 3 Nano 30B A3B output price moved from $0.20/M to $0.24/M (20% up). (details)
- Claude Sonnet 5.5 cached input price moved from $0.20/M to $0.10/M (50% down). (details)
- MiniMax-M2.5-highspeed cached input price moved from $0.06/M to $0.03/M (50% down). (details)
- Qwen-VL OCR input price moved from $0.72/M to $0.07/M (90% down). (details)
- Qwen-VL OCR output price moved from $0.72/M to $0.16/M (78% down). (details)
New benchmark results
- PRBench Finance: 35 new results, including GPT-6 Astra, Claude Fable 5, GPT-5.6 Sol… (board)
- AudioMC: 20 new results, including Gemini 3.8 Flash, Inkling Small, Gemini 2.5 Flash… (board)
- AA Intelligence Index v4.3.2: 5 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
- Humanity's Last Exam (AA): 5 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
- AA-LCR: 5 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
- AA-Omniscience: 5 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
- CritPt: 5 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
- AA-Omniscience: accuracy: 5 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
- AA-Omniscience: non-hallucination: 5 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
- SciCode (AA): 4 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
- GDPval-AA v2.1: 4 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
- Analyst Agent (AA): 4 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
Score revisions
- Claude Sonnet 5.5 (xhigh) on LMArena WebDev: 1786 → 1773. (details)
- MiMo-V2.6-Pro on LMArena WebDev: 1618 → 1630. (details)
- GLM 5.3 Flash on LMArena WebDev: 1616 → 1610. (details)
- Qwen3.8-Flash-Next on LMArena WebDev: 1638 → 1632. (details)
- Mistral Large 4 (high) on LiveBench: 70.5% → 71.8%. (details)
- Mistral Large 4 (high) on LiveBench Reasoning: 80.9% → 83.9%. (details)
- Mistral Large 4 (high) on LiveBench Coding: 75.4% → 77.2%. (details)
- Mistral Large 4 (high) on LiveBench Data Analysis: 74.9% → 76.5%. (details)
- Mistral Large 4 (high) on LiveBench Language: 47% → 49.6%. (details)
- Step 5 Preview (high) on LMArena WebDev: 1570 → 1564. (details)
Speed changes
- Claude Opus 5 time to first answer changed from 22.25 s to 11.69 s. (details)
- GPT-5.6 Sol time to first answer changed from 22.53 s to 12.84 s. (details)
- GPT-5.5 Pro output speed changed from 26 tok/s to 12 tok/s. (details)
- Muse Spark 1.3 output speed changed from 219 tok/s to 368 tok/s. (details)
- Muse Spark 1.3 thinking time changed from 9.12 s to 5.43 s. (details)
- GLM 5.3 time to first token changed from 0.73 s to 1.08 s. (details)
- Gemini 3.8 Flash output speed changed from 216 tok/s to 125 tok/s. (details)
- Claude Opus 4.7 time to first token changed from 4.59 s to 1.61 s. (details)
- Qwen3.8 2.4T A95B time to first token changed from 5.22 s to 2.26 s. (details)
- Muse Spark 1.1 time to first token changed from 1.54 s to 2.28 s. (details)
- Claude Sonnet 5.5 time to first answer changed from 7.16 s to 2.05 s. (details)
- Claude Sonnet 5.5 response time changed from 11.94 s to 6.96 s. (details)
- Claude Opus 4.7 time to first answer changed from 0.94 s to 1.50 s. (details)
- GPT-5.1 output speed changed from 138 tok/s to 85 tok/s. (details)
- GPT-5.6 Luna time to first answer changed from 15.88 s to 9.06 s. (details)
- …and 36 more.
Today's top five
- Claude Opus 5.5 — 72.0
- Claude Fable 5.1 — 70.8
- GPT-6 Astra — 70.2
- Claude Opus 5.5 — 70.2
- Gemini 4 Argon — 70.0
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.