BenchLeader

Daily digest, 12 Sept 2026

10 new models, 9 top-10 rank changes, 6 price changes, 389 new or updated benchmark results across 20 sources.

BenchLeader refreshed 1263 models from 20 of 20 sources. 10 new models, 9 top-10 rank changes, 6 price changes, 389 new or updated benchmark results across 20 sources.

New models

  • Granite 3.3 8B (IBM) appeared with 3 benchmark results. (details)
  • Granite 4.0 Small (IBM) appeared with 3 benchmark results. (details)
  • Marin 8B (Unknown) appeared with 3 benchmark results. (details)
  • Mistral v0 3 7B (Mistral AI) appeared with 3 benchmark results. (details)
  • Olmo 2 13B November 2024 (Ai2) appeared with 3 benchmark results. (details)
  • Olmo 2 32B March 2025 (Ai2) appeared with 3 benchmark results. (details)
  • Olmo 2 7B November 2024 (Ai2) appeared with 3 benchmark results. (details)
  • Olmoe 1B 7B January 2025 (Ai2) appeared with 3 benchmark results. (details)
  • Palmyra Fin (Writer) appeared with 3 benchmark results. (details)
  • Palmyra Med (Writer) appeared with 3 benchmark results. (details)

Movement in the top 10

  • Claude Fable 5.1 moved from #7 to #1 in the BenchLeader Index. (details)
  • GPT-6 Astra (max) moved from #1 to #2 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (xhigh) moved from #4 to #3 in the BenchLeader Index. (details)
  • Claude Fable 5 moved from #2 to #4 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (max) moved from #3 to #5 in the BenchLeader Index. (details)
  • GPT-6 Astra (high) moved from #5 to #6 in the BenchLeader Index. (details)
  • GPT-6 Astra moved from #8 to #7 in the BenchLeader Index. (details)
  • Claude Fable 5.1 (high) moved from #9 to #8 in the BenchLeader Index. (details)
  • Claude Opus 5 moved from #6 to #9 in the BenchLeader Index. (details)

Price changes

  • Qwen3 235B A22B 2507 input price moved from $0.22/M to $0.20/M (9% down). (details)
  • Qwen3 235B A22B 2507 output price moved from $0.88/M to $0.60/M (32% down). (details)
  • Qwen3 235B A22B 2507 (thinking) input price moved from $0.22/M to $0.20/M (9% down). (details)
  • Qwen3 235B A22B 2507 (thinking) output price moved from $0.88/M to $0.60/M (32% down). (details)
  • Llama 3.1 70B input price moved from $0.56/M to $0.72/M (29% up). (details)
  • Llama 3.1 70B output price moved from $0.56/M to $0.72/M (29% up). (details)

New benchmark results

  • HELM Capabilities mean: 58 new results, including Gemini 3 Pro, o3, GPT-5… (board)
  • MMLU-Pro (HELM): 58 new results, including Gemini 3 Pro, o3, GPT-5… (board)
  • GPQA Diamond (HELM): 58 new results, including Gemini 3 Pro, o3, GPT-5… (board)
  • IFEval (HELM): 58 new results, including Gemini 3 Pro, o3, GPT-5… (board)
  • WildBench (HELM): 58 new results, including Gemini 3 Pro, o3, GPT-5… (board)
  • Omni-MATH (HELM): 58 new results, including Gemini 3 Pro, o3, GPT-5… (board)
  • IOI: 20 new results, including Claude Fable 5.1, GPT-6 Astra, Muse Spark 1.3… (board)
  • LMArena Text: 2 new results, including GPT-6 Astra, Grok 4.6 (board)
  • LMArena Hard Prompts: 2 new results, including GPT-6 Astra, Grok 4.6 (board)
  • LMArena Coding: 2 new results, including GPT-6 Astra, Grok 4.6 (board)
  • LMArena Agent: 1 new result, including Muse Spark 1.3 (board)
  • LMArena WebDev: 1 new result, including DeepSeek V4.1 Flash (board)

Score revisions

  • Claude Fable 5.1 (max) on LMArena WebDev: 1764 → 1758. (details)
  • Claude Fable 5.1 (max) on LMArena Agent: 14.5 → 13.9. (details)
  • Claude Opus 5 on IOI: 91.7% → 84.3%. (details)
  • GPT-5.6 Sol (max) on IOI: 86.7% → 91.2%. (details)
  • Qwen3 8 (max) on IOI: 73% → 68.9%. (details)
  • GLM 5.3 Flash on LMArena Coding: 1534 → 1524. (details)
  • Grok 4.6 (high) on LMArena WebDev: 1624 → 1618. (details)
  • Kimi K3 (low) on AA Intelligence Index: 34.5 → 30.5. (details)
  • Qwen3.8 27B on LMArena Coding: 1494 → 1500. (details)
  • GPT-5.6 Luna (max) on IOI: 72.9% → 61.8%. (details)
  • Mistral Medium 3.5 on LMArena Agent: -9.3 → -10.1. (details)
  • GPT-5.3-Codex (xhigh) on IOI: 43.8% → 53.8%. (details)
  • Nemotron 3.5 Lightning 30B A3b Nvfp4 on LMArena Coding: 1432 → 1424. (details)

Speed changes

  • Claude Opus 4.7 time to first token changed from 4.55 s to 1.39 s. (details)
  • GPT-5.5 time to first answer changed from 6.18 s to 4.01 s. (details)
  • Muse Spark 1.1 time to first token changed from 1.51 s to 2.57 s. (details)
  • GPT-5.4 Pro output speed changed from 2 tok/s to 1 tok/s. (details)
  • GPT-5.4 Pro time to first token changed from 65.61 s to 6.51 s. (details)
  • GLM 5.3 Flash time to first token changed from 1.14 s to 1.70 s. (details)
  • GPT-5.2-Codex time to first token changed from 3.55 s to 5.93 s. (details)
  • DeepSeek V4 Flash time to first token changed from 0.87 s to 1.26 s. (details)
  • GLM 4.7 time to first token changed from 0.80 s to 1.71 s. (details)
  • DeepSeek V3.1 output speed changed from 18 tok/s to 29 tok/s. (details)
  • Inkling Small time to first token changed from 0.86 s to 0.43 s. (details)
  • Claude Sonnet 4.5 time to first token changed from 1.10 s to 1.56 s. (details)
  • Nemotron 3 Super time to first token changed from 2.94 s to 1.05 s. (details)
  • Qwen3.6 27B time to first token changed from 3.25 s to 2.11 s. (details)
  • Kimi K2 time to first token changed from 1.52 s to 0.94 s. (details)
  • …and 38 more.

Today's top five

  1. Claude Fable 5.1 — 70.7
  2. GPT-6 Astra — 70.5
  3. Claude Fable 5.1 — 70.4
  4. Claude Fable 5 — 70.2
  5. Claude Fable 5.1 — 70.1

This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.

Get this in your inbox

No ads, no tracking, unsubscribe in one click.
What to receive