BenchLeader

Daily digest, 22 Sept 2026

3 new models, 3 top-10 rank changes, 40 price changes, 283 new or updated benchmark results across 22 sources.

BenchLeader refreshed 1539 models from 22 of 22 sources. 3 new models, 3 top-10 rank changes, 40 price changes, 283 new or updated benchmark results across 22 sources.

New models

  • MiMo-V2.6-Pro (Xiaomi) appeared with 4 benchmark results. (details)
  • GPT-6 Astra (OpenAI) appeared with 1 benchmark result. (details)
  • SWE-2 (max) (Cognition) appeared with 1 benchmark result. (details)

Movement in the top 10

  • GPT-6 Astra (max) moved from #2 to #1 in the BenchLeader Index. (details)
  • GPT-6 Astra (high) moved from #3 to #2 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (high) moved from #1 to #3 in the BenchLeader Index. (details)

Price changes

  • Grok 4 input price moved from $1.25/M to $3.00/M (140% up). (details)
  • Grok 4 output price moved from $2.50/M to $15.00/M (500% up). (details)
  • Qwen3.5 Plus 2026-04-20 (thinking) input price moved from $0.29/M to $0.12/M (60% down). (details)
  • Qwen3.5 Plus 2026-04-20 (thinking) output price moved from $1.72/M to $0.69/M (60% down). (details)
  • Grok 4.1 (thinking) input price moved from $1.25/M to $0.20/M (84% down). (details)
  • Grok 4.1 (thinking) output price moved from $2.50/M to $0.50/M (80% down). (details)
  • Qwen3 Max (max) input price moved from $1.29/M to $0.36/M (72% down). (details)
  • Qwen3 Max (max) output price moved from $7.75/M to $1.43/M (81% down). (details)
  • Grok 4 Fast (thinking) input price moved from $1.25/M to $0.20/M (84% down). (details)
  • Grok 4 Fast (thinking) output price moved from $2.50/M to $0.50/M (80% down). (details)
  • Qwen3.5 Flash input price moved from $0.17/M to $0.03/M (83% down). (details)
  • Qwen3.5 Flash output price moved from $1.03/M to $0.29/M (72% down). (details)
  • Grok 3 input price moved from $1.25/M to $4.00/M (220% up). (details)
  • Grok 3 output price moved from $2.50/M to $20.00/M (700% up). (details)
  • Nemotron 3 Super (thinking) cost per task moved from 1.1 to 1.6 (55% up). (details)
  • Grok 3 mini (thinking) input price moved from $1.25/M to $0.30/M (76% down). (details)
  • Grok 3 mini (thinking) output price moved from $2.50/M to $0.50/M (80% down). (details)
  • Nemotron 3 Nano 30B A3B (thinking) input price moved from $0.06/M to $0.05/M (17% down). (details)
  • Nemotron 3 Nano 30B A3B (thinking) output price moved from $0.24/M to $0.20/M (17% down). (details)
  • Nemotron 3 Nano Omni 30B A3b (thinking) output price moved from $0.80/M to $1.10/M (38% up). (details)
  • Grok 4.1 (no reasoning) input price moved from $1.25/M to $0.20/M (84% down). (details)
  • Grok 4.1 (no reasoning) output price moved from $2.50/M to $0.50/M (80% down). (details)
  • Nemotron 3 Nano 30B A3B (no reasoning) input price moved from $0.06/M to $0.05/M (17% down). (details)
  • Nemotron 3 Nano 30B A3B (no reasoning) output price moved from $0.24/M to $0.20/M (17% down). (details)
  • Grok 4 Fast (no reasoning) input price moved from $1.25/M to $0.20/M (84% down). (details)
  • Grok 4 Fast (no reasoning) output price moved from $2.50/M to $0.50/M (80% down). (details)
  • Nemotron 3 Nano 30B A3B input price moved from $0.06/M to $0.05/M (17% down). (details)
  • Nemotron 3 Nano 30B A3B output price moved from $0.24/M to $0.20/M (17% down). (details)
  • Qwen3.5 Plus 2026-04-20 input price moved from $0.29/M to $0.12/M (60% down). (details)
  • Qwen3.5 Plus 2026-04-20 output price moved from $1.72/M to $0.69/M (60% down). (details)
  • …and 10 more.

New benchmark results

  • LMCA: 48 new results, including GPT-6 Astra, Claude Fable 5.1, Claude Fable 5.1… (board)
  • DTBench: 48 new results, including GPT-6 Astra, Claude Fable 5.1, Claude Fable 5.1… (board)
  • APEX-Agents: 13 new results, including GPT-5.6 Terra, GLM 5.3 Flash, Gemini 3.8 Flash… (board)
  • GDP.pdf: 6 new results, including Claude Fable 5.1, Claude Fable 5.1, Muse Spark 1.3… (board)
  • SciCode: 6 new results, including Muse Spark 1.3, Muse Spark 1.3, Step 5 Preview… (board)
  • BALROG: 5 new results, including GPT-6 Astra, Claude Opus 5, GPT-5.6 Sol… (board)
  • CursorBench: 5 new results, including Muse Spark 1.3, Muse Spark 1.3, Grok 4.7… (board)
  • BTF-3: 5 new results, including Muse Spark 1.3, GPT-6 Astra, GPT-5.6 Sol… (board)
  • FrontierCode: 5 new results, including GLM 5.3, Gemini 3.8 Flash, GLM 5.3 Flash… (board)
  • ALE-Bench: 4 new results, including GPT-6 Astra, Gemini 3.8 Flash, DeepSeek V4.1 Flash… (board)
  • FrontierSWE: 4 new results, including GPT-6 Astra, Claude Opus 5, Claude Fable 5… (board)
  • CyberBench: 4 new results, including GPT-6 Astra, Muse Spark 1.3, Claude Fable 5.1… (board)

Score revisions

  • Claude Fable 5.1 (high) on APEX-Agents: 44.4% → 59.7%. (details)
  • Claude Fable 5.1 (high) on CursorBench: 69.4% → 49.2%. (details)
  • Claude Fable 5.1 (xhigh) on SciCode: 60.1% → 60.9%. (details)
  • Claude Fable 5.1 (xhigh) on CursorBench: 72.8% → 51.6%. (details)
  • Claude Opus 5 (high) on CursorBench: 66.7% → 44.7%. (details)
  • Claude Opus 5 (xhigh) on CursorBench: 69.3% → 46.1%. (details)
  • Claude Opus 5 (max) on SciCode: 55.7% → 56.4%. (details)
  • Claude Opus 5 (max) on APEX-Agents: 43.5% → 65.8%. (details)
  • Claude Opus 5 (max) on CursorBench: 70% → 46.6%. (details)
  • Claude Fable 5.1 (max) on SciCode: 62% → 63.1%. (details)
  • Claude Fable 5.1 (max) on CursorBench: 73.4% → 51.8%. (details)
  • Claude Fable 5.1 (medium) on CursorBench: 68% → 46.8%. (details)
  • GPT-5.6 Sol (max) on SciCode: 56.1% → 57.1%. (details)
  • GPT-5.6 Sol (max) on CursorBench: 67.2% → 41.7%. (details)
  • GPT-5.6 Sol (xhigh) on CursorBench: 64.5% → 37.7%. (details)
  • GPT-5.6 Sol (high) on CursorBench: 63.5% → 35.7%. (details)
  • Claude Fable 5 on APEX-Agents: 45% → 63.6%. (details)
  • GPT-6 Astra on APEX-Agents: 46.7% → 64.7%. (details)
  • Claude Fable 5.1 (low) on CursorBench: 66.2% → 45.1%. (details)
  • Kimi K3 (max) on SciCode: 58.7% → 59.5%. (details)
  • …and 76 more.

Speed changes

  • GLM 5.3 time to first token changed from 2.17 s to 1.27 s. (details)
  • Gemini 3.1 Pro time to first answer changed from 63.62 s to 37.84 s. (details)
  • Gemini 3.1 Pro response time changed from 67.64 s to 41.95 s. (details)
  • Grok 4.7 time to first token changed from 1.41 s to 2.08 s. (details)
  • Grok 4.3 time to first token changed from 0.65 s to 1.69 s. (details)
  • GPT-5.1-Codex output speed changed from 25 tok/s to 40 tok/s. (details)
  • GLM 5V Turbo time to first token changed from 3.96 s to 5.95 s. (details)
  • Ling 3.0 Flash VL time to first token changed from 0.92 s to 1.53 s. (details)
  • Nemotron 3 Super thinking time changed from 8.72 s to 12.00 s. (details)
  • Nemotron 3.5 Lightning time to first token changed from 0.30 s to 0.43 s. (details)
  • Qwen3 30B A3B 2507 time to first token changed from 0.81 s to 1.20 s. (details)
  • Ling Flash 2.0 response time changed from 70.56 s to 97.75 s. (details)
  • GLM 4.6V time to first token changed from 5.42 s to 2.81 s. (details)
  • Gemma 3 12B output speed changed from 7 tok/s to 11 tok/s. (details)
  • Mercury 2.5 output speed changed from 56 tok/s to 186 tok/s. (details)

Today's top five

  1. GPT-6 Astra — 71.9
  2. GPT-6 Astra — 71.7
  3. Claude Fable 5.1 — 71.4
  4. Claude Fable 5.1 — 71.4
  5. Claude Fable 5.1 — 71.3

This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.

Get this in your inbox

No ads, no tracking, unsubscribe in one click.
What to receive