Daily digest, 5 Oct 2026
8 top-10 rank changes, 6 price changes, 65 new or updated benchmark results across 23 sources.
BenchLeader refreshed 1550 models from 23 of 23 sources. 8 top-10 rank changes, 6 price changes, 65 new or updated benchmark results across 23 sources.
Movement in the top 10
- Claude Opus 5.5 (high) moved from #7 to #2 in the BenchLeader Index. (details)
- Claude Opus 5.5 (thinking) moved from #2 to #3 in the BenchLeader Index. (details)
- Claude Opus 5.5 (max) moved from #3 to #4 in the BenchLeader Index. (details)
- Claude Fable 5.1 (xhigh) moved from #4 to #6 in the BenchLeader Index. (details)
- Gemini 4 Argon (high) moved from #11 to #7 in the BenchLeader Index. (details)
- Claude Fable 5.1 (high) moved from #6 to #8 in the BenchLeader Index. (details)
- Claude Opus 5.5 (xhigh) moved from #8 to #10 in the BenchLeader Index. (details)
- GPT-6.1 Sol (max) moved from #10 to #11 in the BenchLeader Index. (details)
Price changes
- GPT-5.6 Sol Pro (max) input price moved from $4.00/M to $2.00/M (50% down). (details)
- GPT-5.6 Sol Pro (max) output price moved from $20.00/M to $10.00/M (50% down). (details)
- GPT-5.6 Sol Pro (xhigh) input price moved from $4.00/M to $2.00/M (50% down). (details)
- GPT-5.6 Sol Pro (xhigh) output price moved from $20.00/M to $10.00/M (50% down). (details)
- GPT-5.6 Sol Pro input price moved from $4.00/M to $2.00/M (50% down). (details)
- GPT-5.6 Sol Pro output price moved from $20.00/M to $10.00/M (50% down). (details)
New benchmark results
- SciCode: 18 new results, including Claude Opus 5.5, Gemini 4 Argon, Claude Opus 5.5… (board)
- APEX-Agents: 11 new results, including Claude Opus 5.5, Qwen3.8 Max, MiMo-V2.6-Pro… (board)
- CursorBench: 9 new results, including Claude Opus 5.5, Claude Opus 5.5, Claude Opus 5.5… (board)
- ProofBench: 5 new results, including Muse Spark 1.3, MiMo-V2.6-Pro, Step 5 Preview… (board)
- GSO-Bench: 4 new results, including GPT-6 Astra, Claude Fable 5, Claude Fable 5.1… (board)
- GDP.pdf: 3 new results, including GPT-6 Astra, DeepSeek V4.1 Flash, Grok 4.7 (board)
- ALE-Bench: 2 new results, including Claude Opus 5.5, MiMo-V2.6-Pro (board)
- FrontierSWE: 2 new results, including Gemini 4 Argon, GLM 5.3 Flash (board)
- Vending-Bench 2: 2 new results, including Claude Opus 5.5, Grok 4.7 (board)
- Blueprint-Bench 2: 2 new results, including Claude Opus 5.5, Grok 4.7 (board)
- WeirdML: 1 new result, including Gemini 3.8 Flash (board)
- LMCA: 1 new result, including Grok 4.7 (board)
Score revisions
- Claude Sonnet 5 (high) on SciCode: 48.6% → 54.3%. (details)
- Qwen3.8 Max (xhigh) on FrontierSWE: 15.8% → 17.8%. (details)
Speed changes
- GPT-5.5 Pro output speed changed from 5 tok/s to 9 tok/s. (details)
- Muse Spark 1.1 time to first token changed from 2.02 s to 3.55 s. (details)
- GPT-5.4 Pro output speed changed from 2 tok/s to 4 tok/s. (details)
- Claude Sonnet 5 output speed changed from 52 tok/s to 75 tok/s. (details)
- GPT-5 nano time to first token changed from 1.54 s to 2.35 s. (details)
- Ling 3.0 Flash VL output speed changed from 48 tok/s to 80 tok/s. (details)
- Qwen3.5-35B-A3B time to first token changed from 0.76 s to 1.37 s. (details)
- Qwen3.5-27B output speed changed from 16 tok/s to 28 tok/s. (details)
- Qwen3.5 Flash output speed changed from 49 tok/s to 74 tok/s. (details)
- GLM 4.7 Flash output speed changed from 25 tok/s to 44 tok/s. (details)
- GLM 4.5V output speed changed from 22 tok/s to 31 tok/s. (details)
- Qwen3.5-9B time to first token changed from 0.74 s to 1.04 s. (details)
- Gemini 2.5 Flash-Lite output speed changed from 81 tok/s to 137 tok/s. (details)
- MiniMax-M1 output speed changed from 21 tok/s to 32 tok/s. (details)
- Qwen3.6 35B A3B output speed changed from 47 tok/s to 78 tok/s. (details)
- …and 16 more.
Today's top five
- Claude Opus 5.5 — 72.0
- Claude Fable 5.1 — 70.8
- GPT-6 Astra — 70.2
- Claude Opus 5.5 — 70.2
- Gemini 4 Argon — 70.0
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.