BenchLeader

Daily digest, 19 Sept 2026

1 new model, 12 top-10 rank changes, 37 price changes, 6876 new or updated benchmark results across 20 sources.

BenchLeader refreshed 1517 models from 20 of 20 sources. 1 new model, 12 top-10 rank changes, 37 price changes, 6876 new or updated benchmark results across 20 sources.

New models

  • Step 5 Preview (StepFun) appeared with 5 benchmark results, entering the BenchLeader Index at #39. (details)

Movement in the top 10

  • Claude Fable 5.1 (high) moved from #8 to #1 in the BenchLeader Index. (details)
  • GPT-6 Astra (high) moved from #3 to #2 in the BenchLeader Index. (details)
  • GPT-6 Astra (max) moved from #7 to #3 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (xhigh) moved from #10 to #4 in the BenchLeader Index. (details)
  • GPT-6 Astra (xhigh) moved from #9 to #6 in the BenchLeader Index. (details)
  • Claude Opus 5 (high) moved from #11 to #8 in the BenchLeader Index. (details)
  • Claude Opus 5 (xhigh) moved from #20 to #9 in the BenchLeader Index. (details)
  • Claude Opus 5 (max) moved from #16 to #10 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (max) moved from #5 to #12 in the BenchLeader Index. (details)
  • Claude Fable 5 moved from #2 to #16 in the BenchLeader Index. (details)
  • Claude Opus 5 moved from #6 to #27 in the BenchLeader Index. (details)
  • Claude Fable 5.1 moved from #1 to #48 in the BenchLeader Index. (details)

Price changes

  • Qwen3.8 27B (xhigh) input price moved from $0.50/M to $0.40/M (20% down). (details)
  • Qwen3.8 27B (xhigh) output price moved from $3.00/M to $2.55/M (15% down). (details)
  • Gemma 4 26B A4B input price moved from $0.12/M to $0.13/M (8% up). (details)
  • Qwen3.8 27B input price moved from $0.50/M to $0.40/M (20% down). (details)
  • Qwen3.8 27B output price moved from $3.00/M to $2.55/M (15% down). (details)
  • Muse Glimmer input price moved from $0.35/M to $0.30/M (14% down). (details)
  • Muse Glimmer output price moved from $1.50/M to $1.10/M (27% down). (details)
  • Qwen3.8 27B (medium) input price moved from $0.50/M to $0.40/M (20% down). (details)
  • Qwen3.8 27B (medium) output price moved from $3.00/M to $2.55/M (15% down). (details)
  • Qwen3.8 27B (low) input price moved from $0.50/M to $0.40/M (20% down). (details)
  • Qwen3.8 27B (low) output price moved from $3.00/M to $2.55/M (15% down). (details)
  • Grok 4 Fast (thinking) input price moved from $0.20/M to $1.25/M (525% up). (details)
  • Grok 4 Fast (thinking) output price moved from $0.50/M to $2.50/M (400% up). (details)
  • Qwen3.8 27B (no reasoning) input price moved from $0.50/M to $0.40/M (20% down). (details)
  • Qwen3.8 27B (no reasoning) output price moved from $3.00/M to $2.55/M (15% down). (details)
  • Gemma 4 26B A4B (no reasoning) input price moved from $0.12/M to $0.13/M (8% up). (details)
  • nemotron-3-nano-30b-a3b (thinking) input price moved from $0.05/M to $0.06/M (20% up). (details)
  • nemotron-3-nano-30b-a3b (thinking) output price moved from $0.20/M to $0.24/M (20% up). (details)
  • Qwen3.5-9B (no reasoning) input price moved from $0.14/M to $0.10/M (29% down). (details)
  • Qwen3.5-9B (no reasoning) output price moved from $0.20/M to $0.15/M (25% down). (details)
  • Nemotron 3 Nano Omni 30B A3b (thinking) input price moved from $0.09/M to $0.20/M (122% up). (details)
  • Nemotron 3 Nano Omni 30B A3b (thinking) output price moved from $0.36/M to $0.80/M (122% up). (details)
  • gpt-oss-20b (high) input price moved from $0.06/M to $0.05/M (17% down). (details)
  • gpt-oss-20b (low) input price moved from $0.06/M to $0.05/M (17% down). (details)
  • gpt-oss-20b input price moved from $0.06/M to $0.05/M (17% down). (details)
  • Grok 4 Fast (no reasoning) input price moved from $0.20/M to $1.25/M (525% up). (details)
  • Grok 4 Fast (no reasoning) output price moved from $0.50/M to $2.50/M (400% up). (details)
  • Qwen3.5-9B input price moved from $0.14/M to $0.10/M (29% down). (details)
  • Qwen3.5-9B output price moved from $0.20/M to $0.15/M (25% down). (details)
  • nemotron-3-nano-30b-a3b input price moved from $0.05/M to $0.06/M (20% up). (details)
  • …and 7 more.

New benchmark results

  • CritPt: 436 new results, including Claude Fable 5.1, GPT-6 Astra, GPT-6 Astra… (board)
  • LMArena Instruction Following: 366 new results, including GPT-6 Astra, Claude Opus 5, Claude Opus 5… (board)
  • LMArena Creative Writing: 364 new results, including GPT-6 Astra, Claude Opus 5, Claude Opus 5… (board)
  • LMArena Multi-turn: 364 new results, including GPT-6 Astra, Claude Opus 5, Claude Opus 5… (board)
  • LMArena Maths: 351 new results, including Claude Opus 5, Claude Opus 5, Claude Fable 5.1… (board)
  • LMArena Longer Queries: 344 new results, including GPT-6 Astra, Claude Opus 5, Claude Opus 5… (board)
  • GDPval (AA): 219 new results, including Claude Fable 5.1, GPT-6 Astra, GPT-6 Astra… (board)
  • Chess Puzzles: 216 new results, including GPT-6 Astra, Claude Opus 5, Claude Fable 5.1… (board)
  • τ²-Bench Banking (AA): 189 new results, including Claude Fable 5.1, GPT-6 Astra, GPT-6 Astra… (board)
  • DTBench: 148 new results, including Claude Opus 5, GPT-5.6 Sol, GPT-5.5… (board)
  • Mystery Game Puzzles: 121 new results, including GPT-6 Astra, Claude Opus 5, Claude Fable 5.1… (board)
  • LMCA: 116 new results, including Claude Opus 5, GPT-5.6 Sol, GPT-5.5… (board)

Score revisions

  • Qwen3 8 (max) on SciCode (AA): 52.1% → 53.2%. (details)
  • GPT-5 on AA Intelligence Index: 23 → 10.4. (details)
  • GPT-5 on GPQA Diamond (AA): 85.4% → 68.6%. (details)
  • GPT-5 on Humanity's Last Exam (AA): 28.5% → 6.6%. (details)
  • GPT-5 on IFBench: 73.1% → 45%. (details)
  • GPT-5 on AA-LCR: 78.2% → 65%. (details)
  • GPT-5 on τ²-Bench Telecom (AA): 84.8% → 0%. (details)
  • GPT-5 on Terminal-Bench Hard: 32.6% → 12.9%. (details)

Speed changes

  • GPT-5.5 time to first answer changed from 37.96 s to 52.19 s. (details)
  • Qwen3.8 2.4T A95B time to first token changed from 1.05 s to 3.20 s. (details)
  • GPT-5.6 Terra time to first answer changed from 161.88 s to 220.86 s. (details)
  • Gemini 3.8 Flash output speed changed from 293 tok/s to 78 tok/s. (details)
  • Grok 4.6 time to first answer changed from 24.67 s to 33.42 s. (details)
  • Claude Opus 4.7 time to first token changed from 2.83 s to 1.64 s. (details)
  • GPT-5.4 Pro output speed changed from 3 tok/s to 1 tok/s. (details)
  • DeepSeek V4 Pro time to first answer changed from 27.21 s to 55.62 s. (details)
  • Claude Opus 4.8 time to first token changed from 4.06 s to 1.46 s. (details)
  • GPT-5.5 output speed changed from 77 tok/s to 49 tok/s. (details)
  • Claude Sonnet 5 time to first token changed from 1.88 s to 2.72 s. (details)
  • GPT-6 Astra output speed changed from 51 tok/s to 31 tok/s. (details)
  • Gemini 3 Flash time to first token changed from 0.89 s to 1.36 s. (details)
  • Claude Sonnet 5 time to first answer changed from 22.36 s to 30.19 s. (details)
  • Claude Opus 4.6 time to first answer changed from 2.06 s to 19.56 s. (details)
  • …and 85 more.

Today's top five

  1. Claude Fable 5.1 — 72.0
  2. GPT-6 Astra — 71.9
  3. GPT-6 Astra — 71.8
  4. Claude Fable 5.1 — 71.6
  5. Claude Fable 5.1 — 71.3

This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.

Get this in your inbox

No ads, no tracking, unsubscribe in one click.
What to receive