Daily digest, 22 Sept 2026
3 new models, 3 top-10 rank changes, 40 price changes, 283 new or updated benchmark results across 22 sources.
BenchLeader refreshed 1539 models from 22 of 22 sources. 3 new models, 3 top-10 rank changes, 40 price changes, 283 new or updated benchmark results across 22 sources.
New models
- MiMo-V2.6-Pro (Xiaomi) appeared with 4 benchmark results. (details)
- GPT-6 Astra (OpenAI) appeared with 1 benchmark result. (details)
- SWE-2 (max) (Cognition) appeared with 1 benchmark result. (details)
Movement in the top 10
- GPT-6 Astra (max) moved from #2 to #1 in the BenchLeader Index. (details)
- GPT-6 Astra (high) moved from #3 to #2 in the BenchLeader Index. (details)
- Claude Fable 5.1 (high) moved from #1 to #3 in the BenchLeader Index. (details)
Price changes
- Grok 4 input price moved from $1.25/M to $3.00/M (140% up). (details)
- Grok 4 output price moved from $2.50/M to $15.00/M (500% up). (details)
- Qwen3.5 Plus 2026-04-20 (thinking) input price moved from $0.29/M to $0.12/M (60% down). (details)
- Qwen3.5 Plus 2026-04-20 (thinking) output price moved from $1.72/M to $0.69/M (60% down). (details)
- Grok 4.1 (thinking) input price moved from $1.25/M to $0.20/M (84% down). (details)
- Grok 4.1 (thinking) output price moved from $2.50/M to $0.50/M (80% down). (details)
- Qwen3 Max (max) input price moved from $1.29/M to $0.36/M (72% down). (details)
- Qwen3 Max (max) output price moved from $7.75/M to $1.43/M (81% down). (details)
- Grok 4 Fast (thinking) input price moved from $1.25/M to $0.20/M (84% down). (details)
- Grok 4 Fast (thinking) output price moved from $2.50/M to $0.50/M (80% down). (details)
- Qwen3.5 Flash input price moved from $0.17/M to $0.03/M (83% down). (details)
- Qwen3.5 Flash output price moved from $1.03/M to $0.29/M (72% down). (details)
- Grok 3 input price moved from $1.25/M to $4.00/M (220% up). (details)
- Grok 3 output price moved from $2.50/M to $20.00/M (700% up). (details)
- Nemotron 3 Super (thinking) cost per task moved from 1.1 to 1.6 (55% up). (details)
- Grok 3 mini (thinking) input price moved from $1.25/M to $0.30/M (76% down). (details)
- Grok 3 mini (thinking) output price moved from $2.50/M to $0.50/M (80% down). (details)
- Nemotron 3 Nano 30B A3B (thinking) input price moved from $0.06/M to $0.05/M (17% down). (details)
- Nemotron 3 Nano 30B A3B (thinking) output price moved from $0.24/M to $0.20/M (17% down). (details)
- Nemotron 3 Nano Omni 30B A3b (thinking) output price moved from $0.80/M to $1.10/M (38% up). (details)
- Grok 4.1 (no reasoning) input price moved from $1.25/M to $0.20/M (84% down). (details)
- Grok 4.1 (no reasoning) output price moved from $2.50/M to $0.50/M (80% down). (details)
- Nemotron 3 Nano 30B A3B (no reasoning) input price moved from $0.06/M to $0.05/M (17% down). (details)
- Nemotron 3 Nano 30B A3B (no reasoning) output price moved from $0.24/M to $0.20/M (17% down). (details)
- Grok 4 Fast (no reasoning) input price moved from $1.25/M to $0.20/M (84% down). (details)
- Grok 4 Fast (no reasoning) output price moved from $2.50/M to $0.50/M (80% down). (details)
- Nemotron 3 Nano 30B A3B input price moved from $0.06/M to $0.05/M (17% down). (details)
- Nemotron 3 Nano 30B A3B output price moved from $0.24/M to $0.20/M (17% down). (details)
- Qwen3.5 Plus 2026-04-20 input price moved from $0.29/M to $0.12/M (60% down). (details)
- Qwen3.5 Plus 2026-04-20 output price moved from $1.72/M to $0.69/M (60% down). (details)
- …and 10 more.
New benchmark results
- LMCA: 48 new results, including GPT-6 Astra, Claude Fable 5.1, Claude Fable 5.1… (board)
- DTBench: 48 new results, including GPT-6 Astra, Claude Fable 5.1, Claude Fable 5.1… (board)
- APEX-Agents: 13 new results, including GPT-5.6 Terra, GLM 5.3 Flash, Gemini 3.8 Flash… (board)
- GDP.pdf: 6 new results, including Claude Fable 5.1, Claude Fable 5.1, Muse Spark 1.3… (board)
- SciCode: 6 new results, including Muse Spark 1.3, Muse Spark 1.3, Step 5 Preview… (board)
- BALROG: 5 new results, including GPT-6 Astra, Claude Opus 5, GPT-5.6 Sol… (board)
- CursorBench: 5 new results, including Muse Spark 1.3, Muse Spark 1.3, Grok 4.7… (board)
- BTF-3: 5 new results, including Muse Spark 1.3, GPT-6 Astra, GPT-5.6 Sol… (board)
- FrontierCode: 5 new results, including GLM 5.3, Gemini 3.8 Flash, GLM 5.3 Flash… (board)
- ALE-Bench: 4 new results, including GPT-6 Astra, Gemini 3.8 Flash, DeepSeek V4.1 Flash… (board)
- FrontierSWE: 4 new results, including GPT-6 Astra, Claude Opus 5, Claude Fable 5… (board)
- CyberBench: 4 new results, including GPT-6 Astra, Muse Spark 1.3, Claude Fable 5.1… (board)
Score revisions
- Claude Fable 5.1 (high) on APEX-Agents: 44.4% → 59.7%. (details)
- Claude Fable 5.1 (high) on CursorBench: 69.4% → 49.2%. (details)
- Claude Fable 5.1 (xhigh) on SciCode: 60.1% → 60.9%. (details)
- Claude Fable 5.1 (xhigh) on CursorBench: 72.8% → 51.6%. (details)
- Claude Opus 5 (high) on CursorBench: 66.7% → 44.7%. (details)
- Claude Opus 5 (xhigh) on CursorBench: 69.3% → 46.1%. (details)
- Claude Opus 5 (max) on SciCode: 55.7% → 56.4%. (details)
- Claude Opus 5 (max) on APEX-Agents: 43.5% → 65.8%. (details)
- Claude Opus 5 (max) on CursorBench: 70% → 46.6%. (details)
- Claude Fable 5.1 (max) on SciCode: 62% → 63.1%. (details)
- Claude Fable 5.1 (max) on CursorBench: 73.4% → 51.8%. (details)
- Claude Fable 5.1 (medium) on CursorBench: 68% → 46.8%. (details)
- GPT-5.6 Sol (max) on SciCode: 56.1% → 57.1%. (details)
- GPT-5.6 Sol (max) on CursorBench: 67.2% → 41.7%. (details)
- GPT-5.6 Sol (xhigh) on CursorBench: 64.5% → 37.7%. (details)
- GPT-5.6 Sol (high) on CursorBench: 63.5% → 35.7%. (details)
- Claude Fable 5 on APEX-Agents: 45% → 63.6%. (details)
- GPT-6 Astra on APEX-Agents: 46.7% → 64.7%. (details)
- Claude Fable 5.1 (low) on CursorBench: 66.2% → 45.1%. (details)
- Kimi K3 (max) on SciCode: 58.7% → 59.5%. (details)
- …and 76 more.
Speed changes
- GLM 5.3 time to first token changed from 2.17 s to 1.27 s. (details)
- Gemini 3.1 Pro time to first answer changed from 63.62 s to 37.84 s. (details)
- Gemini 3.1 Pro response time changed from 67.64 s to 41.95 s. (details)
- Grok 4.7 time to first token changed from 1.41 s to 2.08 s. (details)
- Grok 4.3 time to first token changed from 0.65 s to 1.69 s. (details)
- GPT-5.1-Codex output speed changed from 25 tok/s to 40 tok/s. (details)
- GLM 5V Turbo time to first token changed from 3.96 s to 5.95 s. (details)
- Ling 3.0 Flash VL time to first token changed from 0.92 s to 1.53 s. (details)
- Nemotron 3 Super thinking time changed from 8.72 s to 12.00 s. (details)
- Nemotron 3.5 Lightning time to first token changed from 0.30 s to 0.43 s. (details)
- Qwen3 30B A3B 2507 time to first token changed from 0.81 s to 1.20 s. (details)
- Ling Flash 2.0 response time changed from 70.56 s to 97.75 s. (details)
- GLM 4.6V time to first token changed from 5.42 s to 2.81 s. (details)
- Gemma 3 12B output speed changed from 7 tok/s to 11 tok/s. (details)
- Mercury 2.5 output speed changed from 56 tok/s to 186 tok/s. (details)
Today's top five
- GPT-6 Astra — 71.9
- GPT-6 Astra — 71.7
- Claude Fable 5.1 — 71.4
- Claude Fable 5.1 — 71.4
- Claude Fable 5.1 — 71.3
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.