BenchLeader

Daily digest, 8 Oct 2026

5 new models, 11 top-10 rank changes, 26 price changes, 157 new or updated benchmark results across 23 sources.

BenchLeader refreshed 1557 models from 23 of 23 sources. 5 new models, 11 top-10 rank changes, 26 price changes, 157 new or updated benchmark results across 23 sources.

New models

  • Claude Haiku 5.5 (xhigh) (Anthropic) appeared with 5 benchmark results, entering the BenchLeader Index at #104. (details)
  • Claude Haiku 5.5 (max) (Anthropic) appeared with 5 benchmark results, entering the BenchLeader Index at #121. (details)
  • Claude Haiku 5.5 (low) (Anthropic) appeared with 4 benchmark results. (details)
  • Claude Haiku 5.5 (medium) (Anthropic) appeared with 4 benchmark results. (details)
  • Claude Haiku 5.5 (high) (Anthropic) appeared with 4 benchmark results. (details)

Movement in the top 10

  • Claude Opus 5.5 (max) moved from #10 to #1 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (max) moved from #14 to #2 in the BenchLeader Index. (details)
  • Claude Opus 5.5 (high) moved from #5 to #4 in the BenchLeader Index. (details)
  • GPT-6 Astra (high) moved from #6 to #5 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (xhigh) moved from #7 to #6 in the BenchLeader Index. (details)
  • Gemini 4 Argon (high) moved from #8 to #7 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (high) moved from #9 to #8 in the BenchLeader Index. (details)
  • Claude Opus 5.5 (xhigh) moved from #11 to #9 in the BenchLeader Index. (details)
  • GPT-6.1 Sol (max) moved from #12 to #10 in the BenchLeader Index. (details)
  • Claude Fable 5 moved from #4 to #18 in the BenchLeader Index. (details)
  • Claude Fable 5.1 moved from #1 to #34 in the BenchLeader Index. (details)

Price changes

  • Claude Opus 5.5 (high) cost per task moved from 6 to 1.8 (70% down). (details)
  • Claude Fable 5.1 (xhigh) cost per task moved from 7.6 to 6 (22% down). (details)
  • Claude Fable 5.1 (high) cost per task moved from 7.6 to 3.9 (49% down). (details)
  • Claude Opus 5.5 (xhigh) cost per task moved from 6 to 3.5 (42% down). (details)
  • Claude Sonnet 5.5 (max) cost per task moved from 7.7 to 5.5 (29% down). (details)
  • Claude Sonnet 5.5 (max) cached input price moved from $0.20/M to $0.10/M (50% down). (details)
  • Claude Fable 5.1 (medium) cost per task moved from 7.6 to 3 (61% down). (details)
  • Claude Sonnet 5.5 (xhigh) cost per task moved from 7.7 to 2 (74% down). (details)
  • Claude Sonnet 5.5 (xhigh) cached input price moved from $0.20/M to $0.10/M (50% down). (details)
  • Claude Opus 5.5 (medium) cost per task moved from 6 to 1.3 (78% down). (details)
  • Claude Sonnet 5.5 (high) cost per task moved from 7.7 to 0.9 (88% down). (details)
  • Claude Sonnet 5.5 (high) cached input price moved from $0.20/M to $0.10/M (50% down). (details)
  • Claude Fable 5.1 (low) cost per task moved from 7.6 to 2.4 (69% down). (details)
  • Claude Opus 5.5 (low) cost per task moved from 6 to 0.6 (91% down). (details)
  • Claude Sonnet 5.5 (medium) cost per task moved from 7.7 to 0.5 (94% down). (details)
  • Claude Sonnet 5.5 (medium) cached input price moved from $0.20/M to $0.10/M (50% down). (details)
  • Claude Sonnet 5.5 (low) cost per task moved from 7.7 to 0.3 (95% down). (details)
  • Claude Sonnet 5.5 (low) cached input price moved from $0.20/M to $0.10/M (50% down). (details)
  • DeepSeek V4 Flash (high) input price moved from $0.11/M to $0.13/M (18% up). (details)
  • DeepSeek V4 Flash (high) output price moved from $0.24/M to $0.28/M (17% up). (details)
  • Nemotron 3 Nano 30B A3B input price moved from $0.05/M to $0.06/M (20% up). (details)
  • Nemotron 3 Nano 30B A3B output price moved from $0.20/M to $0.24/M (20% up). (details)
  • Claude Sonnet 5.5 cached input price moved from $0.20/M to $0.10/M (50% down). (details)
  • MiniMax-M2.5-highspeed cached input price moved from $0.06/M to $0.03/M (50% down). (details)
  • Qwen-VL OCR input price moved from $0.72/M to $0.07/M (90% down). (details)
  • Qwen-VL OCR output price moved from $0.72/M to $0.16/M (78% down). (details)

New benchmark results

  • PRBench Finance: 35 new results, including GPT-6 Astra, Claude Fable 5, GPT-5.6 Sol… (board)
  • AudioMC: 20 new results, including Gemini 3.8 Flash, Inkling Small, Gemini 2.5 Flash… (board)
  • AA Intelligence Index v4.3.2: 5 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
  • Humanity's Last Exam (AA): 5 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
  • AA-LCR: 5 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
  • AA-Omniscience: 5 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
  • CritPt: 5 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
  • AA-Omniscience: accuracy: 5 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
  • AA-Omniscience: non-hallucination: 5 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
  • SciCode (AA): 4 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
  • GDPval-AA v2.1: 4 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)
  • Analyst Agent (AA): 4 new results, including Claude Opus 5.5, Claude Fable 5.1, Claude Fable 5… (board)

Score revisions

  • Claude Sonnet 5.5 (xhigh) on LMArena WebDev: 1786 → 1773. (details)
  • MiMo-V2.6-Pro on LMArena WebDev: 1618 → 1630. (details)
  • GLM 5.3 Flash on LMArena WebDev: 1616 → 1610. (details)
  • Qwen3.8-Flash-Next on LMArena WebDev: 1638 → 1632. (details)
  • Mistral Large 4 (high) on LiveBench: 70.5% → 71.8%. (details)
  • Mistral Large 4 (high) on LiveBench Reasoning: 80.9% → 83.9%. (details)
  • Mistral Large 4 (high) on LiveBench Coding: 75.4% → 77.2%. (details)
  • Mistral Large 4 (high) on LiveBench Data Analysis: 74.9% → 76.5%. (details)
  • Mistral Large 4 (high) on LiveBench Language: 47% → 49.6%. (details)
  • Step 5 Preview (high) on LMArena WebDev: 1570 → 1564. (details)

Speed changes

  • Claude Opus 5 time to first answer changed from 22.25 s to 11.69 s. (details)
  • GPT-5.6 Sol time to first answer changed from 22.53 s to 12.84 s. (details)
  • GPT-5.5 Pro output speed changed from 26 tok/s to 12 tok/s. (details)
  • Muse Spark 1.3 output speed changed from 219 tok/s to 368 tok/s. (details)
  • Muse Spark 1.3 thinking time changed from 9.12 s to 5.43 s. (details)
  • GLM 5.3 time to first token changed from 0.73 s to 1.08 s. (details)
  • Gemini 3.8 Flash output speed changed from 216 tok/s to 125 tok/s. (details)
  • Claude Opus 4.7 time to first token changed from 4.59 s to 1.61 s. (details)
  • Qwen3.8 2.4T A95B time to first token changed from 5.22 s to 2.26 s. (details)
  • Muse Spark 1.1 time to first token changed from 1.54 s to 2.28 s. (details)
  • Claude Sonnet 5.5 time to first answer changed from 7.16 s to 2.05 s. (details)
  • Claude Sonnet 5.5 response time changed from 11.94 s to 6.96 s. (details)
  • Claude Opus 4.7 time to first answer changed from 0.94 s to 1.50 s. (details)
  • GPT-5.1 output speed changed from 138 tok/s to 85 tok/s. (details)
  • GPT-5.6 Luna time to first answer changed from 15.88 s to 9.06 s. (details)
  • …and 36 more.

Today's top five

  1. Claude Opus 5.5 — 72.0
  2. Claude Fable 5.1 — 70.8
  3. GPT-6 Astra — 70.2
  4. Claude Opus 5.5 — 70.2
  5. Gemini 4 Argon — 70.0

This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.

Get this in your inbox

No ads, no tracking, unsubscribe in one click.
What to receive