BenchLeader

Daily digest, 3 Oct 2026

1 new model, 9 top-10 rank changes, 22 price changes, 149 new or updated benchmark results across 23 sources.

BenchLeader refreshed 1550 models from 23 of 23 sources. 1 new model, 9 top-10 rank changes, 22 price changes, 149 new or updated benchmark results across 23 sources.

New models

  • Ling 3.1 Flash (Ant Group) is now listed; no independent results yet. (details)

Movement in the top 10

  • Claude Opus 5.5 (max) moved from #4 to #3 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (xhigh) moved from #5 to #4 in the BenchLeader Index. (details)
  • GPT-6 Astra (high) moved from #3 to #5 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (high) moved from #7 to #6 in the BenchLeader Index. (details)
  • Claude Opus 5.5 (high) moved from #6 to #7 in the BenchLeader Index. (details)
  • Claude Opus 5.5 (xhigh) moved from #9 to #8 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (thinking) moved from #8 to #9 in the BenchLeader Index. (details)
  • GPT-6.1 Sol (max) moved from #11 to #10 in the BenchLeader Index. (details)
  • Gemini 4 Argon (high) moved from #10 to #11 in the BenchLeader Index. (details)

Price changes

  • DeepSeek V4 Pro (max) input price moved from $0.43/M to $0.66/M (52% up). (details)
  • DeepSeek V4 Pro (max) output price moved from $0.87/M to $1.98/M (128% up). (details)
  • DeepSeek V4 Pro (max) cached input price moved from $0.00/M to $0.02/M (507% up). (details)
  • DeepSeek V4 Pro (high) input price moved from $0.43/M to $0.66/M (52% up). (details)
  • DeepSeek V4 Pro (high) output price moved from $0.87/M to $1.98/M (128% up). (details)
  • DeepSeek V4 Pro (high) cached input price moved from $0.00/M to $0.02/M (507% up). (details)
  • MiMo-V2.5-Pro input price moved from $0.52/M to $1.00/M (92% up). (details)
  • MiMo-V2.5-Pro output price moved from $1.04/M to $3.00/M (187% up). (details)
  • DeepSeek V4 Pro input price moved from $0.43/M to $0.66/M (52% up). (details)
  • DeepSeek V4 Pro output price moved from $0.87/M to $1.98/M (128% up). (details)
  • DeepSeek V4 Pro cached input price moved from $0.00/M to $0.02/M (507% up). (details)
  • MiMo-V2.5-Pro (no reasoning) input price moved from $0.52/M to $1.00/M (92% up). (details)
  • MiMo-V2.5-Pro (no reasoning) output price moved from $1.04/M to $3.00/M (187% up). (details)
  • DeepSeek V4 Pro (no reasoning) input price moved from $0.43/M to $0.66/M (52% up). (details)
  • DeepSeek V4 Pro (no reasoning) output price moved from $0.87/M to $1.98/M (128% up). (details)
  • DeepSeek V4 Pro (no reasoning) cached input price moved from $0.00/M to $0.02/M (507% up). (details)
  • DeepSeek V4 Pro (low) input price moved from $0.43/M to $0.66/M (52% up). (details)
  • DeepSeek V4 Pro (low) output price moved from $0.87/M to $1.98/M (128% up). (details)
  • DeepSeek V4 Pro (low) cached input price moved from $0.00/M to $0.02/M (507% up). (details)
  • DeepSeek V4 Pro (xhigh) input price moved from $0.43/M to $0.66/M (52% up). (details)
  • DeepSeek V4 Pro (xhigh) output price moved from $0.87/M to $1.98/M (128% up). (details)
  • DeepSeek V4 Pro (xhigh) cached input price moved from $0.00/M to $0.02/M (507% up). (details)

New benchmark results

  • MultiChallenge: 28 new results, including Muse Spark, Muse Spark 1.1, GPT-5.4 Pro… (board)
  • LMArena Agent: 4 new results, including GPT-6.1 Sol, Step 5 Preview, Claude Sonnet 5.5… (board)
  • LMArena Text: 3 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview (board)
  • LMArena Hard Prompts: 3 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview (board)
  • LMArena Coding: 3 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview (board)
  • LMArena Vision: 3 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview (board)
  • LMArena Creative Writing: 3 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview (board)
  • LMArena Instruction Following: 3 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview (board)
  • LMArena Multi-turn: 3 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview (board)
  • LMArena Longer Queries: 3 new results, including GPT-6.1 Sol, Claude Sonnet 5.5, Step 5 Preview (board)
  • AA Intelligence Index v4.3.2: 3 new results, including Grok 4.7, Claude Sonnet 5.5, Kimi K2.6 (board)
  • SciCode (AA): 3 new results, including Grok 4.7, Kimi K2.6, Gemma 4 E4B (board)

Score revisions

  • Claude Opus 5.5 (thinking) on AA-Briefcase: 1822 → 1808. (details)
  • Claude Fable 5.1 (xhigh) on AA-Briefcase: 1669 → 1657. (details)
  • Claude Fable 5.1 (high) on AA-Briefcase: 1592 → 1581. (details)
  • Claude Opus 5.5 (high) on AA-Briefcase: 1705 → 1690. (details)
  • Claude Opus 5.5 (xhigh) on AA-Briefcase: 1780 → 1768. (details)
  • Claude Opus 5 (high) on AA-Briefcase: 1573 → 1562. (details)
  • Claude Opus 5 (max) on AA-Briefcase: 1673 → 1662. (details)
  • Claude Opus 5 (max) on LMArena Agent: 8.5 → 7.9. (details)
  • Claude Opus 5 (xhigh) on AA-Briefcase: 1649 → 1636. (details)
  • GPT-5.6 Sol (xhigh) on LMArena Agent: 7.1 → 6.5. (details)
  • GPT-6 Sol (max) on GDPval (AA): 49.4% → 50.3%. (details)
  • GPT-6 Sol (max) on LMArena Agent: 10.6 → 9.7. (details)
  • Claude Sonnet 5.5 (high) on GDPval (AA): 50.9% → 52.6%. (details)
  • GPT-5.5 (high) on GDPval (AA): 40.9% → 41.5%. (details)
  • Claude Opus 5 (medium) on AA-Briefcase: 1441 → 1435. (details)
  • GPT-6 Sol (xhigh) on GDPval (AA): 46.8% → 47.8%. (details)
  • GLM 5.3 (max) on GDPval (AA): 57.2% → 57.8%. (details)
  • GPT-6 Sol (high) on MMMU-Pro: 81.2% → 82%. (details)
  • GPT-6 Sol (high) on GDPval (AA): 43.8% → 44.8%. (details)
  • GPT-5.6 Terra (max) on GDPval (AA): 46.6% → 47.7%. (details)
  • …and 37 more.

Speed changes

  • GPT-5.5 time to first answer changed from 15.32 s to 21.23 s. (details)
  • GPT-6 Sol output speed changed from 63 tok/s to 87 tok/s. (details)
  • Grok 4.7 time to first answer changed from 78.69 s to 49.32 s. (details)
  • Grok 4.7 response time changed from 85.56 s to 55.49 s. (details)
  • Muse Spark 1.1 output speed changed from 215 tok/s to 124 tok/s. (details)
  • Claude Sonnet 5 time to first answer changed from 16.43 s to 25.45 s. (details)
  • Kimi K2.6 output speed changed from 69 tok/s to 40 tok/s. (details)
  • Ling 3.0 Flash VL output speed changed from 142 tok/s to 48 tok/s. (details)
  • DeepSeek V3.2 output speed changed from 29 tok/s to 15 tok/s. (details)
  • MiMo-V2.5 time to first answer changed from 58.82 s to 38.10 s. (details)
  • Solar Pro 4 output speed changed from 80 tok/s to 109 tok/s. (details)
  • Grok Build 0.1 output speed changed from 48 tok/s to 65 tok/s. (details)
  • GPT-5 time to first answer changed from 7.89 s to 15.19 s. (details)
  • GPT-5 response time changed from 15.03 s to 21.78 s. (details)
  • Qwen3.5 397B A17B output speed changed from 39 tok/s to 53 tok/s. (details)
  • …and 17 more.

Today's top five

  1. Claude Opus 5.5 — 72.0
  2. Claude Fable 5.1 — 70.8
  3. GPT-6 Astra — 70.2
  4. Claude Opus 5.5 — 70.2
  5. Gemini 4 Argon — 70.0

This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.

Get this in your inbox

No ads, no tracking, unsubscribe in one click.
What to receive