BenchLeader

Daily digest, 15 Sept 2026

1 new model, 2 top-10 rank changes, 5 price changes, 13 new or updated benchmark results across 20 sources.

BenchLeader refreshed 1360 models from 20 of 20 sources. 1 new model, 2 top-10 rank changes, 5 price changes, 13 new or updated benchmark results across 20 sources.

New models

  • Mercury 2.5 (high) (Inception) appeared with 3 benchmark results. (details)

Movement in the top 10

  • Claude Opus 5 moved from #9 to #8 in the BenchLeader Index. (details)
  • GPT-6 Astra (max) moved from #8 to #9 in the BenchLeader Index. (details)

Price changes

  • Qwen3.8 27B cached input price moved from $0.04/M to $0.05/M (25% up). (details)
  • Qwen3.8 27B (low) cached input price moved from $0.04/M to $0.05/M (25% up). (details)
  • Qwen3.8 27B (medium) cached input price moved from $0.04/M to $0.05/M (25% up). (details)
  • Qwen3.8 27B (no reasoning) cached input price moved from $0.04/M to $0.05/M (25% up). (details)
  • Qwen3.8 27B (xhigh) cached input price moved from $0.04/M to $0.05/M (25% up). (details)

New benchmark results

  • SciCode (AA): 2 new results, including GPT-5.6 Sol, GPT-5.6 Terra (board)
  • LMArena Agent: 1 new result, including DeepSeek V4.1 Flash (board)

Score revisions

  • GPT-6 Astra (max) on LMArena Agent: 12.4 → 11.9. (details)
  • Claude Opus 5 (high) on LMArena Agent: 11.1 → 10.5. (details)
  • Grok 4.6 (xhigh) on LMArena Agent: 3.1 → 2.6. (details)
  • Claude Sonnet 5 (high) on LMArena Agent: 5.9 → 5.4. (details)
  • GLM 5.3 Flash on LMArena Agent: 1.9 → 1.3. (details)
  • Kimi K2.6 on AA Intelligence Index: 31.3 → 27.5. (details)
  • GLM 5.3 (max) on LMArena Agent: 2.6 → 3.1. (details)
  • Inkling Small on LMArena Agent: -7.1 → -7.8. (details)
  • Gemini 3.5 Flash Lite on LMArena Agent: -13.7 → -14.2. (details)
  • Ling 3.0 Flash on AA Intelligence Index: 24.9 → 20.6. (details)

Speed changes

  • Claude Sonnet 5 time to first answer changed from 4.78 s to 2.76 s. (details)
  • Mistral Medium 3.5 time to first token changed from 0.47 s to 0.66 s. (details)
  • Mercury 2 time to first token changed from 0.86 s to 0.53 s. (details)
  • Qwen2.5 72B time to first token changed from 1.02 s to 0.66 s. (details)
  • Nemotron Nano 12B v2 VL output speed changed from 109 tok/s to 148 tok/s. (details)
  • Gemma 3 4B time to first token changed from 1.07 s to 1.59 s. (details)
  • Palmyra X5 output speed changed from 48 tok/s to 28 tok/s. (details)
  • Devstral 2 output speed changed from 19 tok/s to 31 tok/s. (details)
  • Lyria 3 Pro output speed changed from 4 tok/s to 2 tok/s. (details)
  • Lyria 3 Pro time to first token changed from 2.12 s to 3.63 s. (details)

Today's top five

  1. Claude Fable 5.1 — 70.5
  2. Claude Fable 5.1 — 70.4
  3. GPT-6 Astra — 70.2
  4. Claude Fable 5 — 70.1
  5. Claude Fable 5.1 — 69.9

This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.

Get this in your inbox

No ads, no tracking, unsubscribe in one click.
What to receive