Daily digest, 29 Sept 2026
4 new models, 9 top-10 rank changes, 20 price changes, 169 new or updated benchmark results across 22 sources.
BenchLeader refreshed 1534 models from 22 of 22 sources. 4 new models, 9 top-10 rank changes, 20 price changes, 169 new or updated benchmark results across 22 sources.
New models
- Claude Sonnet 5.5 (Anthropic) appeared with 2 benchmark results. (details)
- Claude Opus 5.5 Range (Anthropic) is now listed; no independent results yet. (details)
- Claude Sonnet 5 Range (Anthropic) is now listed; no independent results yet. (details)
- Resonant 1 (Unknown) is now listed; no independent results yet. (details)
Movement in the top 10
- GPT-6 Astra (high) moved from #2 to #3 in the BenchLeader Index. (details)
- Claude Opus 5.5 (thinking) moved from #6 to #4 in the BenchLeader Index. (details)
- Claude Fable 5.1 (xhigh) moved from #3 to #5 in the BenchLeader Index. (details)
- Claude Fable 5.1 (thinking) moved from #4 to #6 in the BenchLeader Index. (details)
- Claude Fable 5.1 (high) moved from #5 to #7 in the BenchLeader Index. (details)
- Claude Opus 5.5 (high) moved from #7 to #8 in the BenchLeader Index. (details)
- Claude Opus 5.5 (xhigh) moved from #10 to #9 in the BenchLeader Index. (details)
- GPT-6 Astra (xhigh) moved from #8 to #10 in the BenchLeader Index. (details)
- Claude Fable 5 (thinking) moved from #9 to #11 in the BenchLeader Index. (details)
Price changes
- Qwen3.8 27B (xhigh) input price moved from $0.42/M to $0.45/M (7% up). (details)
- Qwen3.8 27B (xhigh) cached input price moved from $0.05/M to $0.45/M (800% up). (details)
- Qwen3.8 27B input price moved from $0.42/M to $0.45/M (7% up). (details)
- Qwen3.8 27B cached input price moved from $0.05/M to $0.45/M (800% up). (details)
- Qwen3.8 27B (medium) input price moved from $0.42/M to $0.45/M (7% up). (details)
- Qwen3.8 27B (medium) cached input price moved from $0.05/M to $0.45/M (800% up). (details)
- Qwen3.8 27B (low) input price moved from $0.42/M to $0.45/M (7% up). (details)
- Qwen3.8 27B (low) cached input price moved from $0.05/M to $0.45/M (800% up). (details)
- Qwen3.8 27B (no reasoning) input price moved from $0.42/M to $0.45/M (7% up). (details)
- Qwen3.8 27B (no reasoning) cached input price moved from $0.05/M to $0.45/M (800% up). (details)
- Nemotron 3.5 Lightning input price moved from $0.07/M to $0.06/M (14% down). (details)
- R1 Distill Llama 70B input price moved from $0.80/M to $0.70/M (13% down). (details)
- Magistral Small input price moved from $0.50/M to $0.15/M (70% down). (details)
- Magistral Small output price moved from $1.50/M to $0.60/M (60% down). (details)
- GPT-5.6 Sol Pro (max) input price moved from $2.00/M to $4.00/M (100% up). (details)
- GPT-5.6 Sol Pro (max) output price moved from $10.00/M to $20.00/M (100% up). (details)
- GPT-5.6 Sol Pro (xhigh) input price moved from $2.00/M to $4.00/M (100% up). (details)
- GPT-5.6 Sol Pro (xhigh) output price moved from $10.00/M to $20.00/M (100% up). (details)
- GPT-5.6 Sol Pro input price moved from $2.00/M to $4.00/M (100% up). (details)
- GPT-5.6 Sol Pro output price moved from $10.00/M to $20.00/M (100% up). (details)
New benchmark results
- Tax Agent Bench: 30 new results, including Muse Spark 1.3, Claude Fable 5, Claude Opus 4.7… (board)
- CyberBench: 27 new results, including GPT-5.6 Sol, Muse Spark 1.3, GPT-6 Sol… (board)
- SWE Atlas: Codebase QnA: 22 new results, including Claude Fable 5.1, GPT-6 Astra, Claude Opus 5… (board)
- SciCode: 6 new results, including Claude Opus 5.5, Claude Opus 5.5, GPT-6 Sol… (board)
- ARC-AGI-1: 6 new results, including GPT-6 Sol, GPT-6 Sol, GPT-6 Sol… (board)
- ARC-AGI-2: 6 new results, including GPT-6 Sol, GPT-6 Sol, GPT-6 Sol… (board)
- ARC-AGI-3: 6 new results, including GPT-6 Sol, GPT-6 Sol, GPT-6 Sol… (board)
- LMArena Vision: 5 new results, including Claude Fable 5, MiMo-V2.6-Pro, Gemini 3.8 Flash… (board)
- APEX-Agents: 3 new results, including Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna (board)
- LMCA: 3 new results, including Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna (board)
- DTBench: 3 new results, including Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna (board)
- GDP.pdf: 3 new results, including Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna (board)
Score revisions
- GPT-6 Sol (max) on LMArena Agent: 8.2 → 8.8. (details)
- Grok 4.6 (high) on CyberBench: 66% → 67.7%. (details)
- GLM 5.3 Flash on LMArena Vision: 1299 → 1304. (details)
- DeepSeek V4 Pro (high) on LMArena Agent: 3.2 → 1.7. (details)
- Qwen3.5-9B (thinking) on AA Intelligence Index: 13.7 → 11.2. (details)
- Molmo 2 8B on LMArena Vision: 1075 → 1081. (details)
Speed changes
- Claude Fable 5.1 time to first answer changed from 11.53 s to 24.07 s. (details)
- Claude Fable 5.1 response time changed from 20.27 s to 33.47 s. (details)
- Claude Opus 5.5 time to first answer changed from 29.87 s to 55.59 s. (details)
- Claude Opus 5.5 response time changed from 36.05 s to 62.24 s. (details)
- Claude Opus 5 time to first answer changed from 14.32 s to 21.47 s. (details)
- Claude Opus 5 response time changed from 44.36 s to 60.57 s. (details)
- Muse Spark 1.3 output speed changed from 261 tok/s to 162 tok/s. (details)
- Muse Spark 1.3 time to first answer changed from 32.89 s to 50.06 s. (details)
- Muse Spark 1.3 response time changed from 34.81 s to 53.15 s. (details)
- Muse Spark 1.3 thinking time changed from 7.66 s to 12.35 s. (details)
- GPT-6 Sol time to first answer changed from 112.71 s to 179.57 s. (details)
- GPT-6 Sol response time changed from 118.45 s to 185.93 s. (details)
- GPT-5.5 time to first answer changed from 9.16 s to 12.98 s. (details)
- GPT-5.5 response time changed from 8.16 s to 11.26 s. (details)
- Claude Opus 4.7 time to first answer changed from 15.11 s to 22.40 s. (details)
- …and 25 more.
Today's top five
- Claude Opus 5.5 — 72.0
- Claude Fable 5.1 — 70.8
- GPT-6 Astra — 70.2
- Claude Opus 5.5 — 70.2
- Gemini 4 Argon — 70.0
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.