BenchLeader

Daily digest, 9 Oct 2026

2 new models, 3 top-10 rank changes, 3 price changes, 151 new or updated benchmark results across 23 sources.

BenchLeader refreshed 1561 models from 23 of 23 sources. 2 new models, 3 top-10 rank changes, 3 price changes, 151 new or updated benchmark results across 23 sources.

New models

  • GPT-6 Sol (Daybreak Blue, max) (max) (OpenAI) is now listed at $4.00/M blended; no independent results yet. (details)
  • Ling 3.0 Flash Sante (Ant Group) is now listed at $0.06/M blended; no independent results yet. (details)

Movement in the top 10

  • Gemini 4 Argon (high) moved from #7 to #5 in the BenchLeader Index. (details)
  • GPT-6 Astra (high) moved from #5 to #6 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (xhigh) moved from #6 to #7 in the BenchLeader Index. (details)

Price changes

  • Step 5 Preview cost per task moved from 0.7 to 1 (43% up). (details)
  • Step 5 Preview (high) cost per task moved from 0.7 to 1 (43% up). (details)
  • Nano Banana 2.1 output price moved from $7.50/M to $30.00/M (300% up). (details)

New benchmark results

  • AudioMC (text output): 17 new results, including Inkling Small, Gemini 2.5 Flash, Gemini 2.5 Flash… (board)
  • Harvey LAB: 16 new results, including GPT-6 Astra, GPT-6.1 Sol, Muse Spark 1.3… (board)
  • SciPredict: 15 new results, including Gemini 3 Pro, o3, GPT-5.2… (board)
  • AudioMC (audio output): 15 new results, including Qwen3 Omni 30B A3B Instruct, Gemini 2.5 Flash Native Audio 12 2025, Gemini 2.5 Flash Native Audio 12 2025… (board)
  • AA-Briefcase v1.1: 6 new results, including Gemini 3.1 Pro, Claude Haiku 5.5, Claude Haiku 5.5… (board)
  • AutomationBench: 5 new results, including Gemini 3.1 Pro, Claude Haiku 5.5, Claude Haiku 5.5… (board)
  • GDP.pdf: 5 new results, including Gemini 3.1 Pro, Claude Haiku 5.5, Claude Haiku 5.5… (board)
  • LMArena Maths: 4 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview… (board)
  • LMArena Agent: 2 new results, including MiMo-V2.6-Pro, Mistral Large 4 (board)
  • MLCR: 2 new results, including Gemini 3.1 Pro, Claude Haiku 5.5 (board)
  • LMArena WebDev: 2 new results, including Claude Haiku 5.5, Mistral Large 4 (board)
  • EnterpriseOps-Gym: 1 new result, including Gemini 3.1 Pro (board)

Score revisions

  • Claude Opus 5.5 (max) on Harvey LAB: 91.2% → 4.2%. (details)
  • Claude Fable 5.1 (max) on Harvey LAB: 93% → 6.4%. (details)
  • Claude Fable 5.1 (max) on LMArena Agent: 14.3 → 12.7. (details)
  • GPT-6 Astra (max) on LMArena Agent: 12.3 → 13.1. (details)
  • Claude Opus 5.5 (high) on LMArena Coding: 1539 → 1550. (details)
  • Claude Opus 5.5 (high) on LMArena Agent: 13.8 → 14.3. (details)
  • Claude Opus 5.5 (high) on LMArena Maths: 1511 → 1499. (details)
  • Gemini 4 Argon (high) on LMArena Agent: 7.6 → 9.3. (details)
  • GPT-6.1 Sol (max) on LMArena Agent: 11.2 → 11.7. (details)
  • GPT-6.1 Sol (max) on LMArena Multi-turn: 1493 → 1487. (details)
  • Claude Opus 5 (high) on LMArena Agent: 8.7 → 8. (details)
  • Muse Spark 1.3 (max) on LMArena Maths: 1509 → 1502. (details)
  • Claude Sonnet 5.5 (max) on Harvey LAB: 93.1% → 2.8%. (details)
  • Claude Sonnet 5.5 (max) on LMArena Agent: 12.5 → 12. (details)
  • Claude Sonnet 5.5 (xhigh) on LMArena Creative Writing: 1447 → 1463. (details)
  • Claude Sonnet 5.5 (xhigh) on LMArena Instruction Following: 1484 → 1489. (details)
  • Claude Sonnet 5.5 (xhigh) on LMArena Multi-turn: 1473 → 1484. (details)
  • Claude Sonnet 5.5 (xhigh) on LMArena Longer Queries: 1493 → 1503. (details)
  • GPT-6 Sol (max) on LMArena Creative Writing: 1442 → 1436. (details)
  • Kimi K3 (max) on Harvey LAB: 94.6% → 5.3%. (details)
  • …and 32 more.

Speed changes

  • GPT-6 Astra time to first token changed from 9.55 s to 4.00 s. (details)
  • Claude Opus 5 time to first answer changed from 11.69 s to 17.07 s. (details)
  • Claude Opus 5 time to first token changed from 1.42 s to 3.10 s. (details)
  • GPT-5.6 Sol time to first answer changed from 134.89 s to 85.27 s. (details)
  • Claude Fable 5.1 time to first answer changed from 7.60 s to 4.51 s. (details)
  • Claude Opus 5 output speed changed from 46 tok/s to 67 tok/s. (details)
  • GPT-5.5 Pro output speed changed from 28 tok/s to 15 tok/s. (details)
  • MiMo-V2.6-Pro time to first token changed from 9.47 s to 3.44 s. (details)
  • Grok 4.7 time to first answer changed from 87.05 s to 50.24 s. (details)
  • Grok 4.7 response time changed from 93.42 s to 57.99 s. (details)
  • GLM 5.3 time to first token changed from 1.82 s to 1.08 s. (details)
  • Claude Opus 4.7 time to first token changed from 3.13 s to 1.77 s. (details)
  • Qwen3.8 2.4T A95B time to first token changed from 4.02 s to 1.48 s. (details)
  • GPT-5.5 time to first answer changed from 8.29 s to 5.30 s. (details)
  • Muse Spark 1.1 output speed changed from 153 tok/s to 93 tok/s. (details)
  • …and 72 more.

Today's top five

  1. Claude Opus 5.5 — 72.0
  2. Claude Fable 5.1 — 70.8
  3. GPT-6 Astra — 70.2
  4. Claude Opus 5.5 — 70.2
  5. Gemini 4 Argon — 70.0

This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.

Get this in your inbox

No ads, no tracking, unsubscribe in one click.
What to receive