BenchLeader

Daily digest, 23 Sept 2026

15 new models, 4 top-10 rank changes, 1 new or updated benchmark result across 22 sources.

BenchLeader refreshed 1551 models from 22 of 22 sources. 15 new models, 4 top-10 rank changes, 1 new or updated benchmark result across 22 sources.

New models

  • GPT-6 Sol (max) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #22. (details)
  • GPT-6 Sol (xhigh) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #33. (details)
  • GPT-6 Sol (high) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #36. (details)
  • GPT-6 Sol (medium) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #43. (details)
  • GPT-6 Sol (low) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #82. (details)
  • GPT-6 Luna (max) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #92. (details)
  • GPT-6 Luna (xhigh) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #115. (details)
  • GPT-6 Luna (high) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #136. (details)
  • GPT-6 Luna (medium) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #172. (details)
  • GPT-6 Sol (no reasoning) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #249. (details)
  • GPT-6 Luna (low) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #268. (details)
  • GPT-6 Luna (no reasoning) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #428. (details)
  • GPT-6 Luna Pro (OpenAI) is now listed at $0.20/M blended; no independent results yet. (details)
  • GPT-6 Sol Pro (OpenAI) is now listed at $4.00/M blended; no independent results yet. (details)
  • Qwen3.8 Omni Flash (Alibaba) is now listed at $0.23/M blended; no independent results yet. (details)

Movement in the top 10

  • Claude Fable 5.1 (high) moved from #4 to #3 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (xhigh) moved from #3 to #4 in the BenchLeader Index. (details)
  • Claude Opus 5.5 (xhigh) moved from #9 to #8 in the BenchLeader Index. (details)
  • Claude Fable 5 (thinking) moved from #8 to #9 in the BenchLeader Index. (details)

New benchmark results

  • SimpleBench: 1 new result, including GPT-5.6 Sol Pro (board)

Speed changes

  • Claude Opus 5.5 output speed changed from 48 tok/s to 84 tok/s. (details)
  • Muse Spark 1.1 time to first token changed from 4.98 s to 2.79 s. (details)
  • GPT-5.4 Pro output speed changed from 1 tok/s to 2 tok/s. (details)
  • Grok 4.5 time to first token changed from 2.69 s to 1.22 s. (details)
  • Deepseek v4 Flash Vision time to first token changed from 3.27 s to 1.36 s. (details)
  • MiMo-V2.5-Pro time to first token changed from 2.45 s to 3.54 s. (details)
  • Kimi K2.5 output speed changed from 30 tok/s to 43 tok/s. (details)
  • GPT-5.1-Codex time to first token changed from 8.18 s to 4.58 s. (details)
  • Ling 3.0 Flash VL time to first token changed from 1.41 s to 0.73 s. (details)
  • Qwen3.5 397B A17B time to first token changed from 5.37 s to 1.13 s. (details)
  • Qwen3.5-27B time to first token changed from 2.33 s to 0.90 s. (details)
  • Qwen3.5-122B-A10B output speed changed from 18 tok/s to 50 tok/s. (details)
  • Qwen3.5-122B-A10B time to first token changed from 1.73 s to 0.93 s. (details)
  • GPT-5.1-Codex-Mini time to first token changed from 11.60 s to 5.79 s. (details)
  • Qwen3.5-35B-A3B time to first token changed from 0.71 s to 1.10 s. (details)
  • …and 67 more.

Today's top five

  1. GPT-6 Astra — 71.6
  2. GPT-6 Astra — 71.0
  3. Claude Fable 5.1 — 70.8
  4. Claude Fable 5.1 — 70.8
  5. Claude Fable 5.1 — 70.7

This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.

Get this in your inbox

No ads, no tracking, unsubscribe in one click.
What to receive