BenchLeader

Daily digest, 2 Oct 2026

3 new models, 5 top-10 rank changes, 19 price changes, 399 new or updated benchmark results across 23 sources.

BenchLeader refreshed 1547 models from 23 of 23 sources. 3 new models, 5 top-10 rank changes, 19 price changes, 399 new or updated benchmark results across 23 sources.

New models

  • Gemini 4 Argon (high) (Google) appeared with 12 benchmark results, entering the BenchLeader Index at #10. (details)
  • GPT-6.1 Sol (OpenAI) appeared with 1 benchmark result. (details)
  • Mercury Voice (Inception) is now listed at $0.68/M blended; no independent results yet. (details)

Movement in the top 10

  • Claude Opus 5.5 (thinking) moved from #4 to #2 in the BenchLeader Index. (details)
  • Claude Opus 5.5 (max) moved from #2 to #4 in the BenchLeader Index. (details)
  • Claude Opus 5.5 (high) moved from #8 to #6 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (thinking) moved from #6 to #8 in the BenchLeader Index. (details)
  • GPT-6 Astra (xhigh) moved from #10 to #12 in the BenchLeader Index. (details)

Price changes

  • Muse Glimmer input price moved from $0.30/M to $0.35/M (17% up). (details)
  • Muse Glimmer output price moved from $1.20/M to $1.50/M (25% up). (details)
  • Muse Glimmer (high) input price moved from $0.30/M to $0.35/M (17% up). (details)
  • Muse Glimmer (high) output price moved from $1.20/M to $1.50/M (25% up). (details)
  • MiniMax-M1 input price moved from $0.40/M to $0.55/M (38% up). (details)
  • Nemotron 3.5 Lightning output price moved from $0.22/M to $0.20/M (9% down). (details)
  • Nemotron 3.5 Lightning cost per task moved from 0.1 to 0.1 (7% down). (details)
  • gpt-oss-20b (high) input price moved from $0.05/M to $0.07/M (40% up). (details)
  • gpt-oss-20b (high) output price moved from $0.20/M to $0.30/M (50% up). (details)
  • gpt-oss-20b input price moved from $0.05/M to $0.07/M (40% up). (details)
  • gpt-oss-20b output price moved from $0.20/M to $0.30/M (50% up). (details)
  • gpt-oss-20b (low) input price moved from $0.05/M to $0.07/M (40% up). (details)
  • gpt-oss-20b (low) output price moved from $0.20/M to $0.30/M (50% up). (details)
  • Gemma 3n E4B input price moved from $0.06/M to $0.00/M (100% down). (details)
  • Gemma 3n E4B output price moved from $0.12/M to $0.00/M (100% down). (details)
  • Solar Mini4 input price moved from $0.05/M to $0.10/M (100% up). (details)
  • Solar Mini4 output price moved from $0.20/M to $0.40/M (100% up). (details)
  • gpt-oss-20b (medium) input price moved from $0.05/M to $0.07/M (40% up). (details)
  • gpt-oss-20b (medium) output price moved from $0.20/M to $0.30/M (50% up). (details)

New benchmark results

  • MultiNRC: 43 new results, including Muse Spark, Muse Spark 1.1, GPT-5.4 Pro… (board)
  • PRBench Finance: 34 new results, including GPT-5.6 Sol, Claude Fable 5, GPT-6 Astra… (board)
  • SWE Atlas: Test Writing: 22 new results, including Claude Fable 5.1, GPT-6 Astra, Claude Opus 5… (board)
  • SciCode: 8 new results, including GPT-6 Sol, GPT-6 Sol, GPT-6 Sol… (board)
  • GDPval-AA v2.1: 8 new results, including Kimi K3, DeepSeek V4 Flash, Inkling Small… (board)
  • ARC-AGI-1: 6 new results, including Qwen3.8 27B, GLM 5.3 Flash, Qwen3.8 27B… (board)
  • ARC-AGI-2: 6 new results, including Qwen3.8 27B, GLM 5.3 Flash, Qwen3.8 27B… (board)
  • Terminal-Bench 4.0 (AA): 6 new results, including DeepSeek V4 Pro, hypernova-60b, Gemma 4 E4B… (board)
  • Mystery Game Puzzles: 5 new results, including Claude Opus 5.5, GPT-6 Sol, Grok 4.7… (board)
  • Analyst Agent (AA): 4 new results, including claude-opus-5-5@thinking, Qwen3.8 Max, Step 5 Preview… (board)
  • GPQA Diamond: 4 new results, including GPT-6 Sol, Grok 4.7, Claude Sonnet 5.5… (board)
  • OTIS Mock AIME: 4 new results, including GPT-6 Sol, Grok 4.7, Claude Sonnet 5.5… (board)

Score revisions

  • GPT-6 Astra (max) on LMArena Agent: 10.4 → 12.2. (details)
  • GPT-6 Astra (max) on LMArena Creative Writing: 1453 → 1448. (details)
  • GPT-6 Astra (max) on LMArena Multi-turn: 1495 → 1489. (details)
  • GPT-6 Astra (max) on Vals Index: 66.6 → 63.1. (details)
  • GPT-6 Astra (max) on Terminal-Bench 4.0 (Vals): 57.1% → 59.6%. (details)
  • GPT-6 Astra (max) on Terminal-Bench Science: 65.7% → 62.9%. (details)
  • Claude Opus 5.5 (max) on LMArena WebDev: 1827 → 1815. (details)
  • Claude Opus 5.5 (high) on LMArena Hard Prompts: 1541 → 1533. (details)
  • Claude Opus 5.5 (high) on LMArena Coding: 1547 → 1538. (details)
  • Claude Opus 5.5 (high) on LMArena Agent: 11.8 → 13.8. (details)
  • Claude Opus 5.5 (high) on LMArena Creative Writing: 1521 → 1515. (details)
  • Claude Opus 5.5 (high) on LMArena Multi-turn: 1519 → 1497. (details)
  • Claude Sonnet 5.5 (xhigh) on AutomationBench: 64.7% → 65.5%. (details)
  • Claude Sonnet 5.5 (xhigh) on AA-Briefcase: 1746 → 1751. (details)
  • Claude Opus 5 (high) on LMArena Agent: 9.4 → 8.8. (details)
  • Claude Fable 5.1 (max) on LMArena Agent: 14.1 → 14.6. (details)
  • Claude Opus 5 (max) on LMArena Agent: 9.5 → 8.5. (details)
  • Muse Spark 1.3 (max) on LMArena Maths: 1505 → 1512. (details)
  • Muse Spark 1.3 (max) on LMArena Creative Writing: 1450 → 1459. (details)
  • Muse Spark 1.3 (max) on Vals Index: 64.5 → 58.2. (details)
  • …and 173 more.

Speed changes

  • GPT-6 Astra time to first answer changed from 40.69 s to 58.39 s. (details)
  • GPT-6 Astra response time changed from 50.65 s to 69.40 s. (details)
  • Claude Fable 5.1 output speed changed from 27 tok/s to 48 tok/s. (details)
  • GPT-5.6 Sol time to first answer changed from 24.12 s to 41.79 s. (details)
  • GPT-5.6 Sol response time changed from 30.45 s to 48.24 s. (details)
  • GPT-6 Astra output speed changed from 27 tok/s to 38 tok/s. (details)
  • GPT-5.4 time to first token changed from 2.10 s to 1.32 s. (details)
  • GPT-6 Sol time to first answer changed from 11.02 s to 19.20 s. (details)
  • GPT-6 Sol response time changed from 18.05 s to 27.08 s. (details)
  • Claude Opus 4.8 time to first answer changed from 34.44 s to 62.22 s. (details)
  • Claude Opus 4.8 response time changed from 44.07 s to 70.70 s. (details)
  • Grok 4.7 time to first answer changed from 55.26 s to 78.69 s. (details)
  • Grok 4.7 response time changed from 62.14 s to 85.56 s. (details)
  • Grok 4.7 time to first token changed from 1.00 s to 2.30 s. (details)
  • Claude Sonnet 5 time to first answer changed from 6.93 s to 13.99 s. (details)
  • …and 90 more.

Today's top five

  1. Claude Opus 5.5 — 72.0
  2. Claude Fable 5.1 — 70.8
  3. GPT-6 Astra — 70.2
  4. Claude Opus 5.5 — 70.2
  5. Gemini 4 Argon — 70.0

This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.

Get this in your inbox

No ads, no tracking, unsubscribe in one click.
What to receive