Daily digest, 19 Sept 2026
1 new model, 12 top-10 rank changes, 37 price changes, 6876 new or updated benchmark results across 20 sources.
BenchLeader refreshed 1517 models from 20 of 20 sources. 1 new model, 12 top-10 rank changes, 37 price changes, 6876 new or updated benchmark results across 20 sources.
New models
- Step 5 Preview (StepFun) appeared with 5 benchmark results, entering the BenchLeader Index at #39. (details)
Movement in the top 10
- Claude Fable 5.1 (high) moved from #8 to #1 in the BenchLeader Index. (details)
- GPT-6 Astra (high) moved from #3 to #2 in the BenchLeader Index. (details)
- GPT-6 Astra (max) moved from #7 to #3 in the BenchLeader Index. (details)
- Claude Fable 5.1 (xhigh) moved from #10 to #4 in the BenchLeader Index. (details)
- GPT-6 Astra (xhigh) moved from #9 to #6 in the BenchLeader Index. (details)
- Claude Opus 5 (high) moved from #11 to #8 in the BenchLeader Index. (details)
- Claude Opus 5 (xhigh) moved from #20 to #9 in the BenchLeader Index. (details)
- Claude Opus 5 (max) moved from #16 to #10 in the BenchLeader Index. (details)
- Claude Fable 5.1 (max) moved from #5 to #12 in the BenchLeader Index. (details)
- Claude Fable 5 moved from #2 to #16 in the BenchLeader Index. (details)
- Claude Opus 5 moved from #6 to #27 in the BenchLeader Index. (details)
- Claude Fable 5.1 moved from #1 to #48 in the BenchLeader Index. (details)
Price changes
- Qwen3.8 27B (xhigh) input price moved from $0.50/M to $0.40/M (20% down). (details)
- Qwen3.8 27B (xhigh) output price moved from $3.00/M to $2.55/M (15% down). (details)
- Gemma 4 26B A4B input price moved from $0.12/M to $0.13/M (8% up). (details)
- Qwen3.8 27B input price moved from $0.50/M to $0.40/M (20% down). (details)
- Qwen3.8 27B output price moved from $3.00/M to $2.55/M (15% down). (details)
- Muse Glimmer input price moved from $0.35/M to $0.30/M (14% down). (details)
- Muse Glimmer output price moved from $1.50/M to $1.10/M (27% down). (details)
- Qwen3.8 27B (medium) input price moved from $0.50/M to $0.40/M (20% down). (details)
- Qwen3.8 27B (medium) output price moved from $3.00/M to $2.55/M (15% down). (details)
- Qwen3.8 27B (low) input price moved from $0.50/M to $0.40/M (20% down). (details)
- Qwen3.8 27B (low) output price moved from $3.00/M to $2.55/M (15% down). (details)
- Grok 4 Fast (thinking) input price moved from $0.20/M to $1.25/M (525% up). (details)
- Grok 4 Fast (thinking) output price moved from $0.50/M to $2.50/M (400% up). (details)
- Qwen3.8 27B (no reasoning) input price moved from $0.50/M to $0.40/M (20% down). (details)
- Qwen3.8 27B (no reasoning) output price moved from $3.00/M to $2.55/M (15% down). (details)
- Gemma 4 26B A4B (no reasoning) input price moved from $0.12/M to $0.13/M (8% up). (details)
- nemotron-3-nano-30b-a3b (thinking) input price moved from $0.05/M to $0.06/M (20% up). (details)
- nemotron-3-nano-30b-a3b (thinking) output price moved from $0.20/M to $0.24/M (20% up). (details)
- Qwen3.5-9B (no reasoning) input price moved from $0.14/M to $0.10/M (29% down). (details)
- Qwen3.5-9B (no reasoning) output price moved from $0.20/M to $0.15/M (25% down). (details)
- Nemotron 3 Nano Omni 30B A3b (thinking) input price moved from $0.09/M to $0.20/M (122% up). (details)
- Nemotron 3 Nano Omni 30B A3b (thinking) output price moved from $0.36/M to $0.80/M (122% up). (details)
- gpt-oss-20b (high) input price moved from $0.06/M to $0.05/M (17% down). (details)
- gpt-oss-20b (low) input price moved from $0.06/M to $0.05/M (17% down). (details)
- gpt-oss-20b input price moved from $0.06/M to $0.05/M (17% down). (details)
- Grok 4 Fast (no reasoning) input price moved from $0.20/M to $1.25/M (525% up). (details)
- Grok 4 Fast (no reasoning) output price moved from $0.50/M to $2.50/M (400% up). (details)
- Qwen3.5-9B input price moved from $0.14/M to $0.10/M (29% down). (details)
- Qwen3.5-9B output price moved from $0.20/M to $0.15/M (25% down). (details)
- nemotron-3-nano-30b-a3b input price moved from $0.05/M to $0.06/M (20% up). (details)
- …and 7 more.
New benchmark results
- CritPt: 436 new results, including Claude Fable 5.1, GPT-6 Astra, GPT-6 Astra… (board)
- LMArena Instruction Following: 366 new results, including GPT-6 Astra, Claude Opus 5, Claude Opus 5… (board)
- LMArena Creative Writing: 364 new results, including GPT-6 Astra, Claude Opus 5, Claude Opus 5… (board)
- LMArena Multi-turn: 364 new results, including GPT-6 Astra, Claude Opus 5, Claude Opus 5… (board)
- LMArena Maths: 351 new results, including Claude Opus 5, Claude Opus 5, Claude Fable 5.1… (board)
- LMArena Longer Queries: 344 new results, including GPT-6 Astra, Claude Opus 5, Claude Opus 5… (board)
- GDPval (AA): 219 new results, including Claude Fable 5.1, GPT-6 Astra, GPT-6 Astra… (board)
- Chess Puzzles: 216 new results, including GPT-6 Astra, Claude Opus 5, Claude Fable 5.1… (board)
- τ²-Bench Banking (AA): 189 new results, including Claude Fable 5.1, GPT-6 Astra, GPT-6 Astra… (board)
- DTBench: 148 new results, including Claude Opus 5, GPT-5.6 Sol, GPT-5.5… (board)
- Mystery Game Puzzles: 121 new results, including GPT-6 Astra, Claude Opus 5, Claude Fable 5.1… (board)
- LMCA: 116 new results, including Claude Opus 5, GPT-5.6 Sol, GPT-5.5… (board)
Score revisions
- Qwen3 8 (max) on SciCode (AA): 52.1% → 53.2%. (details)
- GPT-5 on AA Intelligence Index: 23 → 10.4. (details)
- GPT-5 on GPQA Diamond (AA): 85.4% → 68.6%. (details)
- GPT-5 on Humanity's Last Exam (AA): 28.5% → 6.6%. (details)
- GPT-5 on IFBench: 73.1% → 45%. (details)
- GPT-5 on AA-LCR: 78.2% → 65%. (details)
- GPT-5 on τ²-Bench Telecom (AA): 84.8% → 0%. (details)
- GPT-5 on Terminal-Bench Hard: 32.6% → 12.9%. (details)
Speed changes
- GPT-5.5 time to first answer changed from 37.96 s to 52.19 s. (details)
- Qwen3.8 2.4T A95B time to first token changed from 1.05 s to 3.20 s. (details)
- GPT-5.6 Terra time to first answer changed from 161.88 s to 220.86 s. (details)
- Gemini 3.8 Flash output speed changed from 293 tok/s to 78 tok/s. (details)
- Grok 4.6 time to first answer changed from 24.67 s to 33.42 s. (details)
- Claude Opus 4.7 time to first token changed from 2.83 s to 1.64 s. (details)
- GPT-5.4 Pro output speed changed from 3 tok/s to 1 tok/s. (details)
- DeepSeek V4 Pro time to first answer changed from 27.21 s to 55.62 s. (details)
- Claude Opus 4.8 time to first token changed from 4.06 s to 1.46 s. (details)
- GPT-5.5 output speed changed from 77 tok/s to 49 tok/s. (details)
- Claude Sonnet 5 time to first token changed from 1.88 s to 2.72 s. (details)
- GPT-6 Astra output speed changed from 51 tok/s to 31 tok/s. (details)
- Gemini 3 Flash time to first token changed from 0.89 s to 1.36 s. (details)
- Claude Sonnet 5 time to first answer changed from 22.36 s to 30.19 s. (details)
- Claude Opus 4.6 time to first answer changed from 2.06 s to 19.56 s. (details)
- …and 85 more.
Today's top five
- Claude Fable 5.1 — 72.0
- GPT-6 Astra — 71.9
- GPT-6 Astra — 71.8
- Claude Fable 5.1 — 71.6
- Claude Fable 5.1 — 71.3
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.