Daily digest, 23 Sept 2026
15 new models, 4 top-10 rank changes, 1 new or updated benchmark result across 22 sources.
BenchLeader refreshed 1551 models from 22 of 22 sources. 15 new models, 4 top-10 rank changes, 1 new or updated benchmark result across 22 sources.
New models
- GPT-6 Sol (max) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #22. (details)
- GPT-6 Sol (xhigh) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #33. (details)
- GPT-6 Sol (high) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #36. (details)
- GPT-6 Sol (medium) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #43. (details)
- GPT-6 Sol (low) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #82. (details)
- GPT-6 Luna (max) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #92. (details)
- GPT-6 Luna (xhigh) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #115. (details)
- GPT-6 Luna (high) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #136. (details)
- GPT-6 Luna (medium) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #172. (details)
- GPT-6 Sol (no reasoning) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #249. (details)
- GPT-6 Luna (low) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #268. (details)
- GPT-6 Luna (no reasoning) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #428. (details)
- GPT-6 Luna Pro (OpenAI) is now listed at $0.20/M blended; no independent results yet. (details)
- GPT-6 Sol Pro (OpenAI) is now listed at $4.00/M blended; no independent results yet. (details)
- Qwen3.8 Omni Flash (Alibaba) is now listed at $0.23/M blended; no independent results yet. (details)
Movement in the top 10
- Claude Fable 5.1 (high) moved from #4 to #3 in the BenchLeader Index. (details)
- Claude Fable 5.1 (xhigh) moved from #3 to #4 in the BenchLeader Index. (details)
- Claude Opus 5.5 (xhigh) moved from #9 to #8 in the BenchLeader Index. (details)
- Claude Fable 5 (thinking) moved from #8 to #9 in the BenchLeader Index. (details)
New benchmark results
- SimpleBench: 1 new result, including GPT-5.6 Sol Pro (board)
Speed changes
- Claude Opus 5.5 output speed changed from 48 tok/s to 84 tok/s. (details)
- Muse Spark 1.1 time to first token changed from 4.98 s to 2.79 s. (details)
- GPT-5.4 Pro output speed changed from 1 tok/s to 2 tok/s. (details)
- Grok 4.5 time to first token changed from 2.69 s to 1.22 s. (details)
- Deepseek v4 Flash Vision time to first token changed from 3.27 s to 1.36 s. (details)
- MiMo-V2.5-Pro time to first token changed from 2.45 s to 3.54 s. (details)
- Kimi K2.5 output speed changed from 30 tok/s to 43 tok/s. (details)
- GPT-5.1-Codex time to first token changed from 8.18 s to 4.58 s. (details)
- Ling 3.0 Flash VL time to first token changed from 1.41 s to 0.73 s. (details)
- Qwen3.5 397B A17B time to first token changed from 5.37 s to 1.13 s. (details)
- Qwen3.5-27B time to first token changed from 2.33 s to 0.90 s. (details)
- Qwen3.5-122B-A10B output speed changed from 18 tok/s to 50 tok/s. (details)
- Qwen3.5-122B-A10B time to first token changed from 1.73 s to 0.93 s. (details)
- GPT-5.1-Codex-Mini time to first token changed from 11.60 s to 5.79 s. (details)
- Qwen3.5-35B-A3B time to first token changed from 0.71 s to 1.10 s. (details)
- …and 67 more.
Today's top five
- GPT-6 Astra — 71.6
- GPT-6 Astra — 71.0
- Claude Fable 5.1 — 70.8
- Claude Fable 5.1 — 70.8
- Claude Fable 5.1 — 70.7
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.