BenchLeader

Daily digest, 9 Sept 2026

261 new models, 10 top-10 rank changes, 122 price changes, 4579 new or updated benchmark results across 19 sources.

BenchLeader refreshed 1211 models from 19 of 19 sources. 261 new models, 10 top-10 rank changes, 122 price changes, 4579 new or updated benchmark results across 19 sources.

New models

  • Claude Opus 4.6 (thinking) (Anthropic) appeared on 9 benchmarks, entering the BenchLeader Index at #59. (details)
  • Qwen3 6 (max) (Alibaba) appeared on 5 benchmarks, entering the BenchLeader Index at #60. (details)
  • Gemini 3 Flash (thinking) (Google) appeared on 6 benchmarks, entering the BenchLeader Index at #66. (details)
  • Grok 4.3 (medium) (xAI) appeared on 6 benchmarks, entering the BenchLeader Index at #81. (details)
  • Claude Opus 4.7 (no reasoning) (Anthropic) appeared on 6 benchmarks, entering the BenchLeader Index at #94. (details)
  • Grok 4.3 (low) (xAI) appeared on 6 benchmarks, entering the BenchLeader Index at #102. (details)
  • KAT-Coder-Pro V2 (KwaiKAT) appeared on 5 benchmarks, entering the BenchLeader Index at #131. (details)
  • Qwen3.5 Plus (thinking) (Alibaba) appeared on 6 benchmarks, entering the BenchLeader Index at #147. (details)
  • Gemini 3 Pro (low) (Google) appeared on 5 benchmarks, entering the BenchLeader Index at #161. (details)
  • Claude Sonnet 4.6 (low) (Anthropic) appeared on 6 benchmarks, entering the BenchLeader Index at #174. (details)
  • Claude Opus 4.6 (no reasoning) (Anthropic) appeared on 6 benchmarks, entering the BenchLeader Index at #179. (details)
  • Kimi K2.6 (no reasoning) (Moonshot AI) appeared on 5 benchmarks, entering the BenchLeader Index at #180. (details)
  • GPT-5.1 (thinking) (OpenAI) appeared on 6 benchmarks, entering the BenchLeader Index at #181. (details)
  • GLM-5 (no reasoning) (Zhipu AI) appeared on 5 benchmarks, entering the BenchLeader Index at #197. (details)
  • Command A+ (Cohere) appeared on 6 benchmarks, entering the BenchLeader Index at #213. (details)
  • Grok 3 Mini Fast (high) (xAI) appeared on 8 benchmarks, entering the BenchLeader Index at #222. (details)
  • Qwen3.5 Omni Plus (Alibaba) appeared on 6 benchmarks, entering the BenchLeader Index at #229. (details)
  • Kimi K2.5 (no reasoning) (Moonshot AI) appeared on 6 benchmarks, entering the BenchLeader Index at #241. (details)
  • Gemini 2.5 Flash 09 (thinking) (Google) appeared on 14 benchmarks, entering the BenchLeader Index at #246. (details)
  • Gemma 4 12B (Google) appeared on 6 benchmarks, entering the BenchLeader Index at #258. (details)
  • JT-35B-Flash (China Mobile) appeared on 5 benchmarks, entering the BenchLeader Index at #259. (details)
  • Qwen3.5 27B (no reasoning) (Alibaba) appeared on 6 benchmarks, entering the BenchLeader Index at #262. (details)
  • Grok 3 mini (thinking) (xAI) appeared on 5 benchmarks, entering the BenchLeader Index at #274. (details)
  • Nemotron Cascade 2 30B A3B (NVIDIA) appeared on 5 benchmarks, entering the BenchLeader Index at #275. (details)
  • K-EXAONE (LG AI Research) appeared on 5 benchmarks, entering the BenchLeader Index at #276. (details)
  • Nova 2.0 Lite (high) (thinking) (Amazon) appeared on 6 benchmarks, entering the BenchLeader Index at #282. (details)
  • MiMo-V2.5-Pro (no reasoning) (Xiaomi) appeared on 5 benchmarks, entering the BenchLeader Index at #284. (details)
  • Gemma-4-31B-IT (no reasoning) (NVIDIA) appeared on 6 benchmarks, entering the BenchLeader Index at #285. (details)
  • GLM-4.7 (no reasoning) (Zhipu AI) appeared on 5 benchmarks, entering the BenchLeader Index at #286. (details)
  • Grok 3 Mini Fast (low) (xAI) appeared on 8 benchmarks, entering the BenchLeader Index at #288. (details)
  • …and 231 more.

Movement in the top 10

  • GPT-6 Astra (max) moved from #2 to #1 in the BenchLeader Index. (details)
  • Claude Fable 5 moved from #6 to #2 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (max) moved from #1 to #4 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (high) moved from #10 to #8 in the BenchLeader Index. (details)
  • GPT-6 Astra (high) moved from #7 to #9 in the BenchLeader Index. (details)
  • Claude Opus 5 (high) moved from #8 to #10 in the BenchLeader Index. (details)
  • Claude Opus 5 (max) moved from #4 to #16 in the BenchLeader Index. (details)
  • Claude Fable 5 (max) moved from #3 to #23 in the BenchLeader Index. (details)
  • Claude Opus 4.7 moved from #9 to #25 in the BenchLeader Index. (details)
  • GPT-5.6 Sol (max) moved from #5 to #26 in the BenchLeader Index. (details)

Price changes

  • GPT-6 Astra (max) input price moved from $10.00/M to $11.00/M (10% up). (details)
  • GPT-6 Astra (max) output price moved from $50.00/M to $55.00/M (10% up). (details)
  • GPT-5.6 Sol (high) input price moved from $4.40/M to $4.00/M (9% down). (details)
  • GPT-5.6 Sol (high) output price moved from $22.00/M to $20.00/M (9% down). (details)
  • GPT-5.6 Sol (xhigh) input price moved from $4.40/M to $4.00/M (9% down). (details)
  • GPT-5.6 Sol (xhigh) output price moved from $22.00/M to $20.00/M (9% down). (details)
  • GPT-5.5 (high) input price moved from $5.50/M to $5.00/M (9% down). (details)
  • GPT-5.5 (high) output price moved from $33.00/M to $30.00/M (9% down). (details)
  • GPT-5.6 Sol (medium) input price moved from $4.40/M to $4.00/M (9% down). (details)
  • GPT-5.6 Sol (medium) output price moved from $22.00/M to $20.00/M (9% down). (details)
  • GPT-5.5 (medium) input price moved from $5.50/M to $5.00/M (9% down). (details)
  • GPT-5.5 (medium) output price moved from $33.00/M to $30.00/M (9% down). (details)
  • GPT-5.6 Sol (low) input price moved from $4.40/M to $4.00/M (9% down). (details)
  • GPT-5.6 Sol (low) output price moved from $22.00/M to $20.00/M (9% down). (details)
  • GPT-6 Astra (no reasoning) input price moved from $10.00/M to $11.00/M (10% up). (details)
  • GPT-6 Astra (no reasoning) output price moved from $50.00/M to $55.00/M (10% up). (details)
  • GPT-5.6 Terra (xhigh) input price moved from $2.20/M to $2.00/M (9% down). (details)
  • GPT-5.6 Terra (xhigh) output price moved from $13.20/M to $12.00/M (9% down). (details)
  • Grok 4.6 (medium) input price moved from $2.20/M to $2.00/M (9% down). (details)
  • Grok 4.6 (medium) output price moved from $6.60/M to $6.00/M (9% down). (details)
  • GPT-5.6 Terra (high) input price moved from $2.20/M to $2.00/M (9% down). (details)
  • GPT-5.6 Terra (high) output price moved from $13.20/M to $12.00/M (9% down). (details)
  • Grok 4.6 (xhigh) input price moved from $2.20/M to $2.00/M (9% down). (details)
  • Grok 4.6 (xhigh) output price moved from $6.60/M to $6.00/M (9% down). (details)
  • GPT-5.5 (low) input price moved from $5.50/M to $5.00/M (9% down). (details)
  • GPT-5.5 (low) output price moved from $33.00/M to $30.00/M (9% down). (details)
  • GPT-5.4 (low) input price moved from $2.75/M to $2.50/M (9% down). (details)
  • GPT-5.4 (low) output price moved from $16.50/M to $15.00/M (9% down). (details)
  • Grok 4.6 (low) input price moved from $2.20/M to $2.00/M (9% down). (details)
  • Grok 4.6 (low) output price moved from $6.60/M to $6.00/M (9% down). (details)
  • …and 92 more.

New benchmark results

  • AA Intelligence Index: 374 new results, including Claude Fable 5, Claude Fable 5.1, Claude Opus 5… (board)
  • GPQA Diamond (AA): 362 new results, including Claude Fable 5, Claude Fable 5.1, Claude Opus 5… (board)
  • Humanity's Last Exam (AA): 360 new results, including Claude Fable 5, Claude Fable 5.1, Claude Opus 5… (board)
  • AA-LCR: 313 new results, including Claude Fable 5, Claude Fable 5.1, Claude Opus 5… (board)
  • AA-Omniscience: 311 new results, including Claude Fable 5, Claude Fable 5.1, Claude Opus 5… (board)
  • IFBench: 260 new results, including Claude Fable 5, GPT-5.6 Sol, GPT-5.6 Sol… (board)
  • τ²-Bench Telecom (AA): 254 new results, including Claude Fable 5, GPT-5.6 Sol, GPT-5.6 Sol… (board)
  • Terminal-Bench Hard: 248 new results, including Claude Fable 5, GPT-5.6 Sol, GPT-5.6 Sol… (board)
  • MMMU-Pro: 183 new results, including Claude Opus 5, GPT-6 Astra, GPT-6 Astra… (board)
  • SciCode (AA): 131 new results, including Claude Fable 5, Claude Fable 5.1, Claude Opus 5… (board)
  • LiveCodeBench: 116 new results, including Claude Fable 5, Claude Opus 5, Claude Fable 5.1… (board)
  • TaxEval: 115 new results, including Claude Fable 5, Claude Opus 5, Claude Fable 5.1… (board)

Score revisions

  • Claude Fable 5.1 (max) on LMArena Agent: 15.9 → 14.5. (details)
  • Claude Opus 5 (high) on LMArena Agent: 12.7 → 11.4. (details)
  • GPT-5.6 Sol (xhigh) on LMArena Agent: 9.3 → 7.7. (details)
  • Claude Opus 5 (max) on LMArena Agent: 11.7 → 10.8. (details)
  • GPT-5.5 on LMArena Agent: 4.4 → 3. (details)
  • Muse Spark 1.1 on LMArena Agent: -1.4 → -2.6. (details)
  • GPT-5.6 Terra (xhigh) on LMArena Agent: 3 → 1.5. (details)
  • Gemini 3.1 Pro on LMArena Agent: -3.7 → -5.3. (details)
  • Qwen3 8 (max) on LMArena Agent: 5.5 → 3.9. (details)
  • Claude Fable 5 (high) on LMArena Agent: 10.2 → 9.2. (details)
  • GPT-5.5 (xhigh) on LMArena Agent: 7.2 → 5.3. (details)
  • Grok 4.5 on LMArena Agent: 5.7 → 3.9. (details)
  • Kimi K3 (max) on LMArena Agent: 8 → 6.6. (details)
  • Grok 4.6 (xhigh) on LMArena Agent: 5 → 3.2. (details)
  • Muse Spark 1.2 (xhigh) on LMArena Agent: 0.5 → -1.6. (details)
  • Claude Opus 4.8 (high) on LMArena Agent: 9.2 → 7.7. (details)
  • GLM-5.3-Flash on LMArena Agent: 3.7 → 2. (details)
  • Claude Sonnet 5 (high) on LMArena Agent: 7.4 → 6.3. (details)
  • Claude Sonnet 4.6 on LMArena Agent: 0.5 → -1. (details)
  • Gemini 3.8 Flash (high) on LMArena Agent: 5.5 → 4. (details)
  • …and 37 more.

Speed changes

  • GPT-6 Astra (max) output speed changed from 28 tok/s to 54 tok/s. (details)
  • GPT-6 Astra (max) time to first token changed from 4.63 s to 7.00 s. (details)
  • Claude Fable 5 output speed changed from 31 tok/s to 63 tok/s. (details)
  • Claude Fable 5.1 (xhigh) output speed changed from 40 tok/s to 58 tok/s. (details)
  • Claude Fable 5.1 (xhigh) time to first token changed from 4.77 s to 6.84 s. (details)
  • Claude Fable 5.1 (max) output speed changed from 40 tok/s to 67 tok/s. (details)
  • Claude Fable 5.1 (max) time to first token changed from 4.77 s to 6.84 s. (details)
  • Claude Fable 5.1 output speed changed from 40 tok/s to 67 tok/s. (details)
  • Claude Fable 5.1 time to first token changed from 4.77 s to 6.84 s. (details)
  • GPT-6 Astra output speed changed from 28 tok/s to 54 tok/s. (details)
  • GPT-6 Astra time to first token changed from 4.63 s to 7.00 s. (details)
  • Claude Fable 5.1 (high) output speed changed from 40 tok/s to 54 tok/s. (details)
  • Claude Fable 5.1 (high) time to first token changed from 4.77 s to 6.84 s. (details)
  • GPT-6 Astra (high) output speed changed from 28 tok/s to 49 tok/s. (details)
  • GPT-6 Astra (high) time to first token changed from 4.63 s to 7.00 s. (details)
  • …and 487 more.

Today's top five

  1. GPT-6 Astra — 70.6
  2. Claude Fable 5 — 70.4
  3. Claude Fable 5.1 — 70.3
  4. Claude Fable 5.1 — 70.2
  5. Claude Opus 5 — 70.0

This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.