Daily digest, 18 Sept 2026
4 new models, 8 top-10 rank changes, 13 new or updated benchmark results across 20 sources.
BenchLeader refreshed 1272 models from 20 of 20 sources. 4 new models, 8 top-10 rank changes, 13 new or updated benchmark results across 20 sources.
New models
- K2 Horizon 0.9B (Institute of Foundation Models) appeared with 3 benchmark results. (details)
- K2 Horizon 3.7B (Institute of Foundation Models) appeared with 3 benchmark results. (details)
- K2 Horizon 7B (Institute of Foundation Models) appeared with 3 benchmark results. (details)
- K2 Horizon MoVA 36B A4B (Institute of Foundation Models) appeared with 3 benchmark results. (details)
Movement in the top 10
- GPT-6 Astra moved from #6 to #4 in the BenchLeader Index. (details)
- Claude Fable 5.1 (max) moved from #7 to #5 in the BenchLeader Index. (details)
- Claude Opus 5 moved from #8 to #6 in the BenchLeader Index. (details)
- GPT-6 Astra (max) moved from #9 to #7 in the BenchLeader Index. (details)
- Claude Fable 5.1 (high) moved from #4 to #8 in the BenchLeader Index. (details)
- GPT-6 Astra (xhigh) moved from #11 to #9 in the BenchLeader Index. (details)
- Claude Fable 5.1 (xhigh) moved from #5 to #10 in the BenchLeader Index. (details)
- Claude Opus 5 (high) moved from #10 to #11 in the BenchLeader Index. (details)
New benchmark results
- Terminal-Bench: 8 new results, including Claude Fable 5.1, Claude Fable 5.1, Claude Opus 5… (board)
- MCP Atlas: 2 new results, including Claude Fable 5.1, Nemotron 3 Ultra (board)
- FrontierMath Tiers 1–3: 1 new result, including Muse Spark 1.3 (board)
- FrontierMath Tier 4: 1 new result, including Muse Spark 1.3 (board)
- OTIS Mock AIME: 1 new result, including Muse Spark 1.3 (board)
Speed changes
- Claude Opus 5 time to first answer changed from 12.18 s to 16.69 s. (details)
- GPT-5.6 Sol time to first answer changed from 29.04 s to 40.15 s. (details)
- Muse Spark 1.1 time to first token changed from 1.87 s to 2.89 s. (details)
- Gemini 3.8 Flash time to first answer changed from 30.49 s to 15.12 s. (details)
- Claude Sonnet 5 time to first answer changed from 4.10 s to 2.63 s. (details)
- DeepSeek V4 Pro 0813 output speed changed from 28 tok/s to 40 tok/s. (details)
- Kimi K2 Thinking time to first token changed from 1.27 s to 1.86 s. (details)
- Mistral Large 3 time to first token changed from 0.59 s to 1.05 s. (details)
- Llama 3.1 Nemotron 70B Instruct output speed changed from 76 tok/s to 36 tok/s. (details)
- Llama 3.1 Nemotron 70B Instruct time to first answer changed from 7.87 s to 11.27 s. (details)
- Llama 3.1 Nemotron 70B Instruct time to first token changed from 7.87 s to 11.27 s. (details)
- Nemotron Nano 12B v2 VL time to first answer changed from 37.62 s to 51.62 s. (details)
- Granite 3.3 8B output speed changed from 16 tok/s to 39 tok/s. (details)
- Granite 3.3 8B time to first answer changed from 26.66 s to 10.35 s. (details)
- Granite 3.3 8B time to first token changed from 26.66 s to 10.35 s. (details)
- …and 6 more.
Today's top five
- Claude Fable 5.1 — 70.5
- Claude Fable 5 — 70.1
- GPT-6 Astra — 70.1
- GPT-6 Astra — 69.8
- Claude Fable 5.1 — 69.6
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.