BenchLeader

Daily digest, 6 Oct 2026

1 new model, 10 top-10 rank changes, 38 price changes, 299 new or updated benchmark results across 23 sources.

BenchLeader refreshed 1547 models from 23 of 23 sources. 1 new model, 10 top-10 rank changes, 38 price changes, 299 new or updated benchmark results across 23 sources.

New models

  • Gemini Nano Banana 2.1 (Google) is now listed at $3.00/M blended; no independent results yet. (details)

Movement in the top 10

  • Claude Fable 5.1 moved from #36 to #1 in the BenchLeader Index. (details)
  • Claude Fable 5 moved from #21 to #3 in the BenchLeader Index. (details)
  • GPT-6 Astra (max) moved from #1 to #4 in the BenchLeader Index. (details)
  • Claude Opus 5.5 (high) moved from #2 to #5 in the BenchLeader Index. (details)
  • GPT-6 Astra (high) moved from #5 to #6 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (xhigh) moved from #6 to #7 in the BenchLeader Index. (details)
  • Gemini 4 Argon (high) moved from #7 to #8 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (high) moved from #8 to #9 in the BenchLeader Index. (details)
  • Claude Opus 5.5 (max) moved from #4 to #10 in the BenchLeader Index. (details)
  • Claude Opus 5.5 (xhigh) moved from #10 to #11 in the BenchLeader Index. (details)

Price changes

  • Claude Opus 5.5 (high) cost per task moved from 1.8 to 6 (228% up). (details)
  • Claude Fable 5.1 (xhigh) cost per task moved from 6 to 7.6 (28% up). (details)
  • Claude Fable 5.1 (high) cost per task moved from 3.9 to 7.6 (95% up). (details)
  • Claude Opus 5.5 (xhigh) cost per task moved from 3.5 to 6 (73% up). (details)
  • Claude Fable 5.1 (medium) cost per task moved from 3 to 7.6 (156% up). (details)
  • Claude Sonnet 5.5 (xhigh) cost per task moved from 2.7 to 7.7 (179% up). (details)
  • Claude Opus 5.5 (medium) cost per task moved from 1.3 to 6 (348% up). (details)
  • Claude Sonnet 5.5 (high) cost per task moved from 1.1 to 7.7 (583% up). (details)
  • Claude Fable 5.1 (low) cost per task moved from 2.4 to 7.6 (222% up). (details)
  • Claude Opus 5.5 (low) cost per task moved from 0.6 to 6 (985% up). (details)
  • Grok 4.6 (medium) cost per task moved from 1.5 to 1.1 (24% down). (details)
  • Grok 4.6 (high) cost per task moved from 1.9 to 1.5 (20% down). (details)
  • Gemini 3.5 Flash (high) cost per task moved from 1.6 to 3.7 (136% up). (details)
  • Grok 4.6 (xhigh) cost per task moved from 2.3 to 1.8 (25% down). (details)
  • Gemini 3.1 Pro cost per task moved from 0.7 to 1.3 (92% up). (details)
  • Gemini 3.6 Flash (high) cost per task moved from 0.9 to 1.6 (72% up). (details)
  • Gemini 3.1 Pro (high) cost per task moved from 0.7 to 1.3 (92% up). (details)
  • Grok 4.5 (high) cost per task moved from 1 to 0.9 (15% down). (details)
  • Grok 4.6 (low) cost per task moved from 0.5 to 0.4 (21% down). (details)
  • Claude Sonnet 5.5 (medium) cost per task moved from 0.6 to 7.7 (1201% up). (details)
  • Claude Sonnet 5.5 (low) cost per task moved from 0.4 to 7.7 (1739% up). (details)
  • Muse Glimmer input price moved from $0.35/M to $0.30/M (14% down). (details)
  • Muse Glimmer output price moved from $1.50/M to $1.20/M (20% down). (details)
  • Gemini 3.5 Flash Lite cost per task moved from 0.1 to 0.2 (56% up). (details)
  • Gemini 2.5 Pro cost per task moved from 0.2 to 0.3 (43% up). (details)
  • Gemini 3.1 Flash Lite cost per task moved from 0 to 0.1 (53% up). (details)
  • Gemini 3.1 Flash Lite (high) cost per task moved from 0 to 0.1 (53% up). (details)
  • Gemma 4 31B (no reasoning) input price moved from $0.14/M to $0.15/M (7% up). (details)
  • Gemini 3.5 Flash Lite (high) cost per task moved from 0.1 to 0.2 (56% up). (details)
  • Gemini 3.5 Flash Lite (minimal) cost per task moved from 0.1 to 0.2 (56% up). (details)
  • …and 8 more.

New benchmark results

  • TutorBench: 26 new results, including Muse Spark, GPT-5.4 Pro, Gemini 3.1 Pro… (board)
  • GDPval-AA v2.1: 7 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
  • ARC-AGI-1: 7 new results, including Grok 4.7, DeepSeek V4.1 Flash, Grok 4.7… (board)
  • ARC-AGI-2: 7 new results, including Grok 4.7, DeepSeek V4.1 Flash, Grok 4.7… (board)
  • AA Intelligence Index v4.3.2: 6 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
  • Humanity's Last Exam (AA): 6 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
  • AA-LCR: 6 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
  • AA-Omniscience: 6 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
  • CritPt: 6 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
  • Terminal-Bench 4.0 (AA): 6 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
  • AA-Omniscience: accuracy: 6 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
  • AA-Omniscience: non-hallucination: 6 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)

Score revisions

  • Claude Opus 5.5 (high) on GDPval (AA): 59.6% → 60.3%. (details)
  • Claude Fable 5.1 (xhigh) on GDPval (AA): 61.1% → 61.8%. (details)
  • Gemini 4 Argon (high) on GDPval (AA): 55.6% → 56.3%. (details)
  • Claude Fable 5.1 (high) on GDPval (AA): 55.9% → 56.8%. (details)
  • Claude Opus 5.5 (xhigh) on GDPval (AA): 66% → 66.8%. (details)
  • Claude Opus 5 (high) on GDPval (AA): 54.1% → 54.8%. (details)
  • Claude Opus 5 (max) on GDPval (AA): 60.4% → 61.2%. (details)
  • Muse Spark 1.3 (max) on GDPval (AA): 58.6% → 59.2%. (details)
  • Claude Opus 5 (xhigh) on GDPval (AA): 58.8% → 59.6%. (details)
  • Claude Fable 5.1 (medium) on GDPval (AA): 51.8% → 52.5%. (details)
  • GPT-5.6 Sol (max) on GDPval (AA): 54.4% → 55.6%. (details)
  • GPT-5.6 Sol (xhigh) on GDPval (AA): 52.4% → 53.6%. (details)
  • Claude Opus 5.5 (medium) on GDPval (AA): 53.8% → 54.3%. (details)
  • GPT-5.6 Sol (high) on GDPval (AA): 49% → 50.2%. (details)
  • Muse Spark 1.3 (xhigh) on GDPval (AA): 55.9% → 56.5%. (details)
  • GPT-5.5 (xhigh) on GDPval (AA): 41.8% → 42.7%. (details)
  • Claude Fable 5.1 (low) on GDPval (AA): 47.5% → 48.4%. (details)
  • Claude Opus 5.5 (low) on GDPval (AA): 36.2% → 36.8%. (details)
  • Claude Opus 5 (medium) on GDPval (AA): 48.8% → 49.6%. (details)
  • Muse Spark on GDPval (AA): 24.3% → 25.1%. (details)
  • …and 125 more.

Speed changes

  • Claude Opus 5.5 output speed changed from 68 tok/s to 97 tok/s. (details)
  • GPT-6.1 Sol time to first answer changed from 110.28 s to 151.76 s. (details)
  • Muse Spark 1.3 output speed changed from 140 tok/s to 192 tok/s. (details)
  • Kimi K3 output speed changed from 34 tok/s to 52 tok/s. (details)
  • Kimi K3 time to first token changed from 2.21 s to 1.27 s. (details)
  • Claude Opus 5 time to first answer changed from 5.27 s to 3.21 s. (details)
  • Grok 4.7 time to first answer changed from 33.12 s to 64.48 s. (details)
  • Grok 4.7 response time changed from 39.49 s to 69.67 s. (details)
  • GPT-5.6 Terra time to first answer changed from 176.80 s to 114.78 s. (details)
  • GLM 5.3 Flash time to first token changed from 2.02 s to 1.20 s. (details)
  • Grok 4.5 time to first answer changed from 9.83 s to 17.49 s. (details)
  • Claude Sonnet 5.5 time to first answer changed from 1.23 s to 7.22 s. (details)
  • Claude Sonnet 5.5 response time changed from 6.29 s to 12.00 s. (details)
  • Qwen3.8 27B time to first token changed from 0.76 s to 1.10 s. (details)
  • Kimi K3 time to first answer changed from 61.39 s to 38.70 s. (details)
  • …and 98 more.

Today's top five

  1. Claude Opus 5.5 — 72.0
  2. Claude Fable 5.1 — 70.8
  3. GPT-6 Astra — 70.2
  4. Claude Opus 5.5 — 70.2
  5. Gemini 4 Argon — 70.0

This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.

Get this in your inbox

No ads, no tracking, unsubscribe in one click.
What to receive