Daily digest, 10 Oct 2026
14 price changes, 50 new or updated benchmark results across 23 sources.
BenchLeader refreshed 1560 models from 23 of 23 sources. 14 price changes, 50 new or updated benchmark results across 23 sources.
Price changes
- Qwen3.8 27B input price moved from $0.45/M to $0.42/M (6% down). (details)
- DeepSeek V4 Flash 0731 (high) cached input price moved from $0.03/M to $0.03/M (7% down). (details)
- DeepSeek V4 Flash 0731 (max) cached input price moved from $0.03/M to $0.03/M (7% down). (details)
- DeepSeek V4 Flash 0731 cached input price moved from $0.03/M to $0.03/M (7% down). (details)
- LongCat 2.0 input price moved from $0.30/M to $0.75/M (150% up). (details)
- LongCat 2.0 output price moved from $1.20/M to $2.95/M (146% up). (details)
- LongCat 2.0 cost per task moved from 0.1 to 0.1 (148% up). (details)
- Nemotron 3 Nano 30B A3B input price moved from $0.06/M to $0.05/M (17% down). (details)
- Nemotron 3 Nano 30B A3B output price moved from $0.24/M to $0.20/M (17% down). (details)
- DeepSeek V4 Flash 0731 (no reasoning) cached input price moved from $0.03/M to $0.03/M (7% down). (details)
- DeepSeek V4 Flash 0731 (low) cached input price moved from $0.03/M to $0.03/M (7% down). (details)
- BGE M3 input price moved from $0.00/M to $0.01/M (1000000000% up). (details)
- Ling 3.0 Flash Sante input price moved from $0.04/M to $0.07/M (79% up). (details)
- Ling 3.0 Flash Sante output price moved from $0.12/M to $0.22/M (79% up). (details)
New benchmark results
- PRBench Finance: 35 new results, including Claude Fable 5, GPT-5.6 Sol, GPT-6 Astra… (board)
- OTIS Mock AIME: 3 new results, including DeepSeek V4.1 Flash, Claude Haiku 5.5, Claude Haiku 5.5 (board)
- GPQA Diamond: 2 new results, including DeepSeek V4.1 Flash, Claude Haiku 5.5 (board)
- FrontierMath Tiers 1–3: 2 new results, including DeepSeek V4.1 Flash, Claude Haiku 5.5 (board)
- FrontierMath Tier 4: 2 new results, including DeepSeek V4.1 Flash, Claude Haiku 5.5 (board)
- Mystery Game Puzzles: 2 new results, including DeepSeek V4.1 Flash, Claude Haiku 5.5 (board)
- SimpleQA Verified: 2 new results, including Claude Haiku 5.5, Mistral Large 4 (board)
Score revisions
- GPT-6.1 Sol (max) on AA-Briefcase v1.1: 1564 → 1557. (details)
- GPT-6.1 Sol (high) on AA-Briefcase v1.1: 1471 → 1465. (details)
Speed changes
- GPT-5.6 Sol time to first token changed from 5.76 s to 3.59 s. (details)
- GPT-5.6 Sol time to first answer changed from 9.70 s to 20.45 s. (details)
- GPT-5.6 Sol response time changed from 17.11 s to 27.03 s. (details)
- GPT-5.5 time to first token changed from 2.23 s to 3.23 s. (details)
- Qwen3.8 2.4T A95B time to first token changed from 0.34 s to 1.48 s. (details)
- GPT-5.5 time to first answer changed from 5.30 s to 7.61 s. (details)
- GLM 5.2 time to first token changed from 0.58 s to 0.89 s. (details)
- GPT-5.4 Pro output speed changed from 40 tok/s to 6 tok/s. (details)
- GPT-5.6 Terra time to first answer changed from 2.64 s to 4.81 s. (details)
- Grok 4.20 time to first token changed from 1.15 s to 0.73 s. (details)
- Gemma 4 26B A4B time to first token changed from 0.87 s to 0.56 s. (details)
- GPT-5 nano time to first token changed from 3.06 s to 1.85 s. (details)
- Muse Glimmer time to first token changed from 1.92 s to 1.12 s. (details)
- Qwen3.5-27B time to first token changed from 1.88 s to 0.91 s. (details)
- Qwen3.5-122B-A10B output speed changed from 45 tok/s to 67 tok/s. (details)
- …and 44 more.
Today's top five
- Claude Opus 5.5 — 72.0
- Claude Fable 5.1 — 70.8
- GPT-6 Astra — 70.2
- Claude Opus 5.5 — 70.2
- Gemini 4 Argon — 70.0
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.