BenchLeader

Daily digest, 11 Oct 2026

91 new or updated benchmark results across 23 sources.

BenchLeader refreshed 1554 models from 23 of 23 sources. 91 new or updated benchmark results across 23 sources.

New benchmark results

  • MultiNRC: 43 new results, including Muse Spark, GPT-5.4 Pro, Muse Spark 1.1… (board)
  • PRBench Legal: 39 new results, including GPT-6 Astra, Claude Fable 5, GPT-5.6 Sol… (board)
  • LMArena Vision: 6 new results, including Claude Opus 5.5, GPT-6 Sol, Muse Spark 1.3… (board)
  • Analyst Agent (AA): 1 new result, including Qwen3.8 27B (board)

Score revisions

  • Step 5 Preview on LMArena Vision: 1278 → 1267. (details)
  • GLM 5.3 Flash on LMArena Vision: 1301 → 1296. (details)

Speed changes

  • Kimi K2 Thinking time to first token changed from 0.97 s to 1.33 s. (details)
  • Qwen3.5-35B-A3B output speed changed from 65 tok/s to 121 tok/s. (details)
  • Qwen3.5-9B output speed changed from 33 tok/s to 53 tok/s. (details)
  • Nemotron 3 Nano 30B A3B output speed changed from 114 tok/s to 198 tok/s. (details)

Today's top five

  1. Claude Opus 5.5 — 72.0
  2. Claude Fable 5.1 — 70.8
  3. GPT-6 Astra — 70.2
  4. Claude Opus 5.5 — 70.2
  5. Gemini 4 Argon — 70.0

This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.

Get this in your inbox

No ads, no tracking, unsubscribe in one click.
What to receive