BenchLeader

Daily digest, 14 Sept 2026

8 top-10 rank changes, 12 new or updated benchmark results across 20 sources.

BenchLeader refreshed 1355 models from 20 of 20 sources. 8 top-10 rank changes, 12 new or updated benchmark results across 20 sources.

Movement in the top 10

  • Claude Fable 5.1 (xhigh) moved from #3 to #2 in the BenchLeader Index. (details)
  • Claude Fable 5 moved from #4 to #3 in the BenchLeader Index. (details)
  • GPT-6 Astra (high) moved from #5 to #4 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (high) moved from #7 to #5 in the BenchLeader Index. (details)
  • GPT-6 Astra moved from #8 to #6 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (max) moved from #6 to #7 in the BenchLeader Index. (details)
  • Claude Opus 5 moved from #9 to #8 in the BenchLeader Index. (details)
  • GPT-6 Astra (max) moved from #2 to #9 in the BenchLeader Index. (details)

New benchmark results

  • LMArena Vision: 4 new results, including Claude Fable 5.1, GPT-6 Astra, Muse Spark 1.3… (board)
  • PRBench Finance: 1 new result, including Muse Spark 1.3 (board)
  • PRBench Legal: 1 new result, including Muse Spark 1.3 (board)
  • Epoch Capabilities Index: 1 new result, including Gemini 3.8 Flash (board)
  • LMArena Text: 1 new result, including Muse Spark 1.3 (board)
  • LMArena Hard Prompts: 1 new result, including Muse Spark 1.3 (board)
  • LMArena Coding: 1 new result, including Muse Spark 1.3 (board)

Score revisions

  • GPT-6 Astra (max) on LMArena Coding: 1537 → 1543. (details)
  • Qwen3.8 27B on LMArena Vision: 1279 → 1272. (details)

Speed changes

  • Gemini 3.8 Flash time to first answer changed from 12.81 s to 17.67 s. (details)
  • Nemotron 3 Ultra 550B A55B time to first token changed from 1.09 s to 0.40 s. (details)
  • Ling Flash 2.0 output speed changed from 3 tok/s to 5 tok/s. (details)
  • Llama 3.1 Nemotron 70B Instruct output speed changed from 60 tok/s to 91 tok/s. (details)
  • Llama 3.1 Nemotron 70B Instruct time to first answer changed from 8.36 s to 4.33 s. (details)
  • Llama 3.1 Nemotron 70B Instruct time to first token changed from 8.36 s to 4.33 s. (details)
  • Nemotron Nano 12B v2 VL output speed changed from 54 tok/s to 82 tok/s. (details)
  • Qwen2.5 Coder 32B Instruct time to first token changed from 0.98 s to 1.40 s. (details)
  • Granite 4.0 H Small output speed changed from 20 tok/s to 12 tok/s. (details)
  • Nemotron 3.5 Lightning time to first token changed from 0.62 s to 0.40 s. (details)
  • Lyria 3 Pro time to first token changed from 3.30 s to 2.12 s. (details)

Today's top five

  1. Claude Fable 5.1 — 70.6
  2. Claude Fable 5.1 — 70.4
  3. Claude Fable 5 — 70.2
  4. GPT-6 Astra — 70.2
  5. Claude Fable 5.1 — 70.0

This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.

Get this in your inbox

No ads, no tracking, unsubscribe in one click.
What to receive