BenchLeader

Daily digest, 18 Sept 2026

4 new models, 8 top-10 rank changes, 13 new or updated benchmark results across 20 sources.

BenchLeader refreshed 1272 models from 20 of 20 sources. 4 new models, 8 top-10 rank changes, 13 new or updated benchmark results across 20 sources.

New models

  • K2 Horizon 0.9B (Institute of Foundation Models) appeared with 3 benchmark results. (details)
  • K2 Horizon 3.7B (Institute of Foundation Models) appeared with 3 benchmark results. (details)
  • K2 Horizon 7B (Institute of Foundation Models) appeared with 3 benchmark results. (details)
  • K2 Horizon MoVA 36B A4B (Institute of Foundation Models) appeared with 3 benchmark results. (details)

Movement in the top 10

  • GPT-6 Astra moved from #6 to #4 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (max) moved from #7 to #5 in the BenchLeader Index. (details)
  • Claude Opus 5 moved from #8 to #6 in the BenchLeader Index. (details)
  • GPT-6 Astra (max) moved from #9 to #7 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (high) moved from #4 to #8 in the BenchLeader Index. (details)
  • GPT-6 Astra (xhigh) moved from #11 to #9 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (xhigh) moved from #5 to #10 in the BenchLeader Index. (details)
  • Claude Opus 5 (high) moved from #10 to #11 in the BenchLeader Index. (details)

New benchmark results

  • Terminal-Bench: 8 new results, including Claude Fable 5.1, Claude Fable 5.1, Claude Opus 5… (board)
  • MCP Atlas: 2 new results, including Claude Fable 5.1, Nemotron 3 Ultra (board)
  • FrontierMath Tiers 1–3: 1 new result, including Muse Spark 1.3 (board)
  • FrontierMath Tier 4: 1 new result, including Muse Spark 1.3 (board)
  • OTIS Mock AIME: 1 new result, including Muse Spark 1.3 (board)

Speed changes

  • Claude Opus 5 time to first answer changed from 12.18 s to 16.69 s. (details)
  • GPT-5.6 Sol time to first answer changed from 29.04 s to 40.15 s. (details)
  • Muse Spark 1.1 time to first token changed from 1.87 s to 2.89 s. (details)
  • Gemini 3.8 Flash time to first answer changed from 30.49 s to 15.12 s. (details)
  • Claude Sonnet 5 time to first answer changed from 4.10 s to 2.63 s. (details)
  • DeepSeek V4 Pro 0813 output speed changed from 28 tok/s to 40 tok/s. (details)
  • Kimi K2 Thinking time to first token changed from 1.27 s to 1.86 s. (details)
  • Mistral Large 3 time to first token changed from 0.59 s to 1.05 s. (details)
  • Llama 3.1 Nemotron 70B Instruct output speed changed from 76 tok/s to 36 tok/s. (details)
  • Llama 3.1 Nemotron 70B Instruct time to first answer changed from 7.87 s to 11.27 s. (details)
  • Llama 3.1 Nemotron 70B Instruct time to first token changed from 7.87 s to 11.27 s. (details)
  • Nemotron Nano 12B v2 VL time to first answer changed from 37.62 s to 51.62 s. (details)
  • Granite 3.3 8B output speed changed from 16 tok/s to 39 tok/s. (details)
  • Granite 3.3 8B time to first answer changed from 26.66 s to 10.35 s. (details)
  • Granite 3.3 8B time to first token changed from 26.66 s to 10.35 s. (details)
  • …and 6 more.

Today's top five

  1. Claude Fable 5.1 — 70.5
  2. Claude Fable 5 — 70.1
  3. GPT-6 Astra — 70.1
  4. GPT-6 Astra — 69.8
  5. Claude Fable 5.1 — 69.6

This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.

Get this in your inbox

No ads, no tracking, unsubscribe in one click.
What to receive