Daily digest, 6 Oct 2026
1 new model, 10 top-10 rank changes, 38 price changes, 299 new or updated benchmark results across 23 sources.
BenchLeader refreshed 1547 models from 23 of 23 sources. 1 new model, 10 top-10 rank changes, 38 price changes, 299 new or updated benchmark results across 23 sources.
New models
- Gemini Nano Banana 2.1 (Google) is now listed at $3.00/M blended; no independent results yet. (details)
Movement in the top 10
- Claude Fable 5.1 moved from #36 to #1 in the BenchLeader Index. (details)
- Claude Fable 5 moved from #21 to #3 in the BenchLeader Index. (details)
- GPT-6 Astra (max) moved from #1 to #4 in the BenchLeader Index. (details)
- Claude Opus 5.5 (high) moved from #2 to #5 in the BenchLeader Index. (details)
- GPT-6 Astra (high) moved from #5 to #6 in the BenchLeader Index. (details)
- Claude Fable 5.1 (xhigh) moved from #6 to #7 in the BenchLeader Index. (details)
- Gemini 4 Argon (high) moved from #7 to #8 in the BenchLeader Index. (details)
- Claude Fable 5.1 (high) moved from #8 to #9 in the BenchLeader Index. (details)
- Claude Opus 5.5 (max) moved from #4 to #10 in the BenchLeader Index. (details)
- Claude Opus 5.5 (xhigh) moved from #10 to #11 in the BenchLeader Index. (details)
Price changes
- Claude Opus 5.5 (high) cost per task moved from 1.8 to 6 (228% up). (details)
- Claude Fable 5.1 (xhigh) cost per task moved from 6 to 7.6 (28% up). (details)
- Claude Fable 5.1 (high) cost per task moved from 3.9 to 7.6 (95% up). (details)
- Claude Opus 5.5 (xhigh) cost per task moved from 3.5 to 6 (73% up). (details)
- Claude Fable 5.1 (medium) cost per task moved from 3 to 7.6 (156% up). (details)
- Claude Sonnet 5.5 (xhigh) cost per task moved from 2.7 to 7.7 (179% up). (details)
- Claude Opus 5.5 (medium) cost per task moved from 1.3 to 6 (348% up). (details)
- Claude Sonnet 5.5 (high) cost per task moved from 1.1 to 7.7 (583% up). (details)
- Claude Fable 5.1 (low) cost per task moved from 2.4 to 7.6 (222% up). (details)
- Claude Opus 5.5 (low) cost per task moved from 0.6 to 6 (985% up). (details)
- Grok 4.6 (medium) cost per task moved from 1.5 to 1.1 (24% down). (details)
- Grok 4.6 (high) cost per task moved from 1.9 to 1.5 (20% down). (details)
- Gemini 3.5 Flash (high) cost per task moved from 1.6 to 3.7 (136% up). (details)
- Grok 4.6 (xhigh) cost per task moved from 2.3 to 1.8 (25% down). (details)
- Gemini 3.1 Pro cost per task moved from 0.7 to 1.3 (92% up). (details)
- Gemini 3.6 Flash (high) cost per task moved from 0.9 to 1.6 (72% up). (details)
- Gemini 3.1 Pro (high) cost per task moved from 0.7 to 1.3 (92% up). (details)
- Grok 4.5 (high) cost per task moved from 1 to 0.9 (15% down). (details)
- Grok 4.6 (low) cost per task moved from 0.5 to 0.4 (21% down). (details)
- Claude Sonnet 5.5 (medium) cost per task moved from 0.6 to 7.7 (1201% up). (details)
- Claude Sonnet 5.5 (low) cost per task moved from 0.4 to 7.7 (1739% up). (details)
- Muse Glimmer input price moved from $0.35/M to $0.30/M (14% down). (details)
- Muse Glimmer output price moved from $1.50/M to $1.20/M (20% down). (details)
- Gemini 3.5 Flash Lite cost per task moved from 0.1 to 0.2 (56% up). (details)
- Gemini 2.5 Pro cost per task moved from 0.2 to 0.3 (43% up). (details)
- Gemini 3.1 Flash Lite cost per task moved from 0 to 0.1 (53% up). (details)
- Gemini 3.1 Flash Lite (high) cost per task moved from 0 to 0.1 (53% up). (details)
- Gemma 4 31B (no reasoning) input price moved from $0.14/M to $0.15/M (7% up). (details)
- Gemini 3.5 Flash Lite (high) cost per task moved from 0.1 to 0.2 (56% up). (details)
- Gemini 3.5 Flash Lite (minimal) cost per task moved from 0.1 to 0.2 (56% up). (details)
- …and 8 more.
New benchmark results
- TutorBench: 26 new results, including Muse Spark, GPT-5.4 Pro, Gemini 3.1 Pro… (board)
- GDPval-AA v2.1: 7 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
- ARC-AGI-1: 7 new results, including Grok 4.7, DeepSeek V4.1 Flash, Grok 4.7… (board)
- ARC-AGI-2: 7 new results, including Grok 4.7, DeepSeek V4.1 Flash, Grok 4.7… (board)
- AA Intelligence Index v4.3.2: 6 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
- Humanity's Last Exam (AA): 6 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
- AA-LCR: 6 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
- AA-Omniscience: 6 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
- CritPt: 6 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
- Terminal-Bench 4.0 (AA): 6 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
- AA-Omniscience: accuracy: 6 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
- AA-Omniscience: non-hallucination: 6 new results, including Claude Fable 5.1, Claude Opus 5.5, Claude Fable 5… (board)
Score revisions
- Claude Opus 5.5 (high) on GDPval (AA): 59.6% → 60.3%. (details)
- Claude Fable 5.1 (xhigh) on GDPval (AA): 61.1% → 61.8%. (details)
- Gemini 4 Argon (high) on GDPval (AA): 55.6% → 56.3%. (details)
- Claude Fable 5.1 (high) on GDPval (AA): 55.9% → 56.8%. (details)
- Claude Opus 5.5 (xhigh) on GDPval (AA): 66% → 66.8%. (details)
- Claude Opus 5 (high) on GDPval (AA): 54.1% → 54.8%. (details)
- Claude Opus 5 (max) on GDPval (AA): 60.4% → 61.2%. (details)
- Muse Spark 1.3 (max) on GDPval (AA): 58.6% → 59.2%. (details)
- Claude Opus 5 (xhigh) on GDPval (AA): 58.8% → 59.6%. (details)
- Claude Fable 5.1 (medium) on GDPval (AA): 51.8% → 52.5%. (details)
- GPT-5.6 Sol (max) on GDPval (AA): 54.4% → 55.6%. (details)
- GPT-5.6 Sol (xhigh) on GDPval (AA): 52.4% → 53.6%. (details)
- Claude Opus 5.5 (medium) on GDPval (AA): 53.8% → 54.3%. (details)
- GPT-5.6 Sol (high) on GDPval (AA): 49% → 50.2%. (details)
- Muse Spark 1.3 (xhigh) on GDPval (AA): 55.9% → 56.5%. (details)
- GPT-5.5 (xhigh) on GDPval (AA): 41.8% → 42.7%. (details)
- Claude Fable 5.1 (low) on GDPval (AA): 47.5% → 48.4%. (details)
- Claude Opus 5.5 (low) on GDPval (AA): 36.2% → 36.8%. (details)
- Claude Opus 5 (medium) on GDPval (AA): 48.8% → 49.6%. (details)
- Muse Spark on GDPval (AA): 24.3% → 25.1%. (details)
- …and 125 more.
Speed changes
- Claude Opus 5.5 output speed changed from 68 tok/s to 97 tok/s. (details)
- GPT-6.1 Sol time to first answer changed from 110.28 s to 151.76 s. (details)
- Muse Spark 1.3 output speed changed from 140 tok/s to 192 tok/s. (details)
- Kimi K3 output speed changed from 34 tok/s to 52 tok/s. (details)
- Kimi K3 time to first token changed from 2.21 s to 1.27 s. (details)
- Claude Opus 5 time to first answer changed from 5.27 s to 3.21 s. (details)
- Grok 4.7 time to first answer changed from 33.12 s to 64.48 s. (details)
- Grok 4.7 response time changed from 39.49 s to 69.67 s. (details)
- GPT-5.6 Terra time to first answer changed from 176.80 s to 114.78 s. (details)
- GLM 5.3 Flash time to first token changed from 2.02 s to 1.20 s. (details)
- Grok 4.5 time to first answer changed from 9.83 s to 17.49 s. (details)
- Claude Sonnet 5.5 time to first answer changed from 1.23 s to 7.22 s. (details)
- Claude Sonnet 5.5 response time changed from 6.29 s to 12.00 s. (details)
- Qwen3.8 27B time to first token changed from 0.76 s to 1.10 s. (details)
- Kimi K3 time to first answer changed from 61.39 s to 38.70 s. (details)
- …and 98 more.
Today's top five
- Claude Opus 5.5 — 72.0
- Claude Fable 5.1 — 70.8
- GPT-6 Astra — 70.2
- Claude Opus 5.5 — 70.2
- Gemini 4 Argon — 70.0
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.