Daily digest, 2 Oct 2026
3 new models, 5 top-10 rank changes, 19 price changes, 399 new or updated benchmark results across 23 sources.
BenchLeader refreshed 1547 models from 23 of 23 sources. 3 new models, 5 top-10 rank changes, 19 price changes, 399 new or updated benchmark results across 23 sources.
New models
- Gemini 4 Argon (high) (Google) appeared with 12 benchmark results, entering the BenchLeader Index at #10. (details)
- GPT-6.1 Sol (OpenAI) appeared with 1 benchmark result. (details)
- Mercury Voice (Inception) is now listed at $0.68/M blended; no independent results yet. (details)
Movement in the top 10
- Claude Opus 5.5 (thinking) moved from #4 to #2 in the BenchLeader Index. (details)
- Claude Opus 5.5 (max) moved from #2 to #4 in the BenchLeader Index. (details)
- Claude Opus 5.5 (high) moved from #8 to #6 in the BenchLeader Index. (details)
- Claude Fable 5.1 (thinking) moved from #6 to #8 in the BenchLeader Index. (details)
- GPT-6 Astra (xhigh) moved from #10 to #12 in the BenchLeader Index. (details)
Price changes
- Muse Glimmer input price moved from $0.30/M to $0.35/M (17% up). (details)
- Muse Glimmer output price moved from $1.20/M to $1.50/M (25% up). (details)
- Muse Glimmer (high) input price moved from $0.30/M to $0.35/M (17% up). (details)
- Muse Glimmer (high) output price moved from $1.20/M to $1.50/M (25% up). (details)
- MiniMax-M1 input price moved from $0.40/M to $0.55/M (38% up). (details)
- Nemotron 3.5 Lightning output price moved from $0.22/M to $0.20/M (9% down). (details)
- Nemotron 3.5 Lightning cost per task moved from 0.1 to 0.1 (7% down). (details)
- gpt-oss-20b (high) input price moved from $0.05/M to $0.07/M (40% up). (details)
- gpt-oss-20b (high) output price moved from $0.20/M to $0.30/M (50% up). (details)
- gpt-oss-20b input price moved from $0.05/M to $0.07/M (40% up). (details)
- gpt-oss-20b output price moved from $0.20/M to $0.30/M (50% up). (details)
- gpt-oss-20b (low) input price moved from $0.05/M to $0.07/M (40% up). (details)
- gpt-oss-20b (low) output price moved from $0.20/M to $0.30/M (50% up). (details)
- Gemma 3n E4B input price moved from $0.06/M to $0.00/M (100% down). (details)
- Gemma 3n E4B output price moved from $0.12/M to $0.00/M (100% down). (details)
- Solar Mini4 input price moved from $0.05/M to $0.10/M (100% up). (details)
- Solar Mini4 output price moved from $0.20/M to $0.40/M (100% up). (details)
- gpt-oss-20b (medium) input price moved from $0.05/M to $0.07/M (40% up). (details)
- gpt-oss-20b (medium) output price moved from $0.20/M to $0.30/M (50% up). (details)
New benchmark results
- MultiNRC: 43 new results, including Muse Spark, Muse Spark 1.1, GPT-5.4 Pro… (board)
- PRBench Finance: 34 new results, including GPT-5.6 Sol, Claude Fable 5, GPT-6 Astra… (board)
- SWE Atlas: Test Writing: 22 new results, including Claude Fable 5.1, GPT-6 Astra, Claude Opus 5… (board)
- SciCode: 8 new results, including GPT-6 Sol, GPT-6 Sol, GPT-6 Sol… (board)
- GDPval-AA v2.1: 8 new results, including Kimi K3, DeepSeek V4 Flash, Inkling Small… (board)
- ARC-AGI-1: 6 new results, including Qwen3.8 27B, GLM 5.3 Flash, Qwen3.8 27B… (board)
- ARC-AGI-2: 6 new results, including Qwen3.8 27B, GLM 5.3 Flash, Qwen3.8 27B… (board)
- Terminal-Bench 4.0 (AA): 6 new results, including DeepSeek V4 Pro, hypernova-60b, Gemma 4 E4B… (board)
- Mystery Game Puzzles: 5 new results, including Claude Opus 5.5, GPT-6 Sol, Grok 4.7… (board)
- Analyst Agent (AA): 4 new results, including claude-opus-5-5@thinking, Qwen3.8 Max, Step 5 Preview… (board)
- GPQA Diamond: 4 new results, including GPT-6 Sol, Grok 4.7, Claude Sonnet 5.5… (board)
- OTIS Mock AIME: 4 new results, including GPT-6 Sol, Grok 4.7, Claude Sonnet 5.5… (board)
Score revisions
- GPT-6 Astra (max) on LMArena Agent: 10.4 → 12.2. (details)
- GPT-6 Astra (max) on LMArena Creative Writing: 1453 → 1448. (details)
- GPT-6 Astra (max) on LMArena Multi-turn: 1495 → 1489. (details)
- GPT-6 Astra (max) on Vals Index: 66.6 → 63.1. (details)
- GPT-6 Astra (max) on Terminal-Bench 4.0 (Vals): 57.1% → 59.6%. (details)
- GPT-6 Astra (max) on Terminal-Bench Science: 65.7% → 62.9%. (details)
- Claude Opus 5.5 (max) on LMArena WebDev: 1827 → 1815. (details)
- Claude Opus 5.5 (high) on LMArena Hard Prompts: 1541 → 1533. (details)
- Claude Opus 5.5 (high) on LMArena Coding: 1547 → 1538. (details)
- Claude Opus 5.5 (high) on LMArena Agent: 11.8 → 13.8. (details)
- Claude Opus 5.5 (high) on LMArena Creative Writing: 1521 → 1515. (details)
- Claude Opus 5.5 (high) on LMArena Multi-turn: 1519 → 1497. (details)
- Claude Sonnet 5.5 (xhigh) on AutomationBench: 64.7% → 65.5%. (details)
- Claude Sonnet 5.5 (xhigh) on AA-Briefcase: 1746 → 1751. (details)
- Claude Opus 5 (high) on LMArena Agent: 9.4 → 8.8. (details)
- Claude Fable 5.1 (max) on LMArena Agent: 14.1 → 14.6. (details)
- Claude Opus 5 (max) on LMArena Agent: 9.5 → 8.5. (details)
- Muse Spark 1.3 (max) on LMArena Maths: 1505 → 1512. (details)
- Muse Spark 1.3 (max) on LMArena Creative Writing: 1450 → 1459. (details)
- Muse Spark 1.3 (max) on Vals Index: 64.5 → 58.2. (details)
- …and 173 more.
Speed changes
- GPT-6 Astra time to first answer changed from 40.69 s to 58.39 s. (details)
- GPT-6 Astra response time changed from 50.65 s to 69.40 s. (details)
- Claude Fable 5.1 output speed changed from 27 tok/s to 48 tok/s. (details)
- GPT-5.6 Sol time to first answer changed from 24.12 s to 41.79 s. (details)
- GPT-5.6 Sol response time changed from 30.45 s to 48.24 s. (details)
- GPT-6 Astra output speed changed from 27 tok/s to 38 tok/s. (details)
- GPT-5.4 time to first token changed from 2.10 s to 1.32 s. (details)
- GPT-6 Sol time to first answer changed from 11.02 s to 19.20 s. (details)
- GPT-6 Sol response time changed from 18.05 s to 27.08 s. (details)
- Claude Opus 4.8 time to first answer changed from 34.44 s to 62.22 s. (details)
- Claude Opus 4.8 response time changed from 44.07 s to 70.70 s. (details)
- Grok 4.7 time to first answer changed from 55.26 s to 78.69 s. (details)
- Grok 4.7 response time changed from 62.14 s to 85.56 s. (details)
- Grok 4.7 time to first token changed from 1.00 s to 2.30 s. (details)
- Claude Sonnet 5 time to first answer changed from 6.93 s to 13.99 s. (details)
- …and 90 more.
Today's top five
- Claude Opus 5.5 — 72.0
- Claude Fable 5.1 — 70.8
- GPT-6 Astra — 70.2
- Claude Opus 5.5 — 70.2
- Gemini 4 Argon — 70.0
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.