Daily digest, 12 Sept 2026
10 new models, 9 top-10 rank changes, 6 price changes, 389 new or updated benchmark results across 20 sources.
BenchLeader refreshed 1263 models from 20 of 20 sources. 10 new models, 9 top-10 rank changes, 6 price changes, 389 new or updated benchmark results across 20 sources.
New models
- Granite 3.3 8B (IBM) appeared with 3 benchmark results. (details)
- Granite 4.0 Small (IBM) appeared with 3 benchmark results. (details)
- Marin 8B (Unknown) appeared with 3 benchmark results. (details)
- Mistral v0 3 7B (Mistral AI) appeared with 3 benchmark results. (details)
- Olmo 2 13B November 2024 (Ai2) appeared with 3 benchmark results. (details)
- Olmo 2 32B March 2025 (Ai2) appeared with 3 benchmark results. (details)
- Olmo 2 7B November 2024 (Ai2) appeared with 3 benchmark results. (details)
- Olmoe 1B 7B January 2025 (Ai2) appeared with 3 benchmark results. (details)
- Palmyra Fin (Writer) appeared with 3 benchmark results. (details)
- Palmyra Med (Writer) appeared with 3 benchmark results. (details)
Movement in the top 10
- Claude Fable 5.1 moved from #7 to #1 in the BenchLeader Index. (details)
- GPT-6 Astra (max) moved from #1 to #2 in the BenchLeader Index. (details)
- Claude Fable 5.1 (xhigh) moved from #4 to #3 in the BenchLeader Index. (details)
- Claude Fable 5 moved from #2 to #4 in the BenchLeader Index. (details)
- Claude Fable 5.1 (max) moved from #3 to #5 in the BenchLeader Index. (details)
- GPT-6 Astra (high) moved from #5 to #6 in the BenchLeader Index. (details)
- GPT-6 Astra moved from #8 to #7 in the BenchLeader Index. (details)
- Claude Fable 5.1 (high) moved from #9 to #8 in the BenchLeader Index. (details)
- Claude Opus 5 moved from #6 to #9 in the BenchLeader Index. (details)
Price changes
- Qwen3 235B A22B 2507 input price moved from $0.22/M to $0.20/M (9% down). (details)
- Qwen3 235B A22B 2507 output price moved from $0.88/M to $0.60/M (32% down). (details)
- Qwen3 235B A22B 2507 (thinking) input price moved from $0.22/M to $0.20/M (9% down). (details)
- Qwen3 235B A22B 2507 (thinking) output price moved from $0.88/M to $0.60/M (32% down). (details)
- Llama 3.1 70B input price moved from $0.56/M to $0.72/M (29% up). (details)
- Llama 3.1 70B output price moved from $0.56/M to $0.72/M (29% up). (details)
New benchmark results
- HELM Capabilities mean: 58 new results, including Gemini 3 Pro, o3, GPT-5… (board)
- MMLU-Pro (HELM): 58 new results, including Gemini 3 Pro, o3, GPT-5… (board)
- GPQA Diamond (HELM): 58 new results, including Gemini 3 Pro, o3, GPT-5… (board)
- IFEval (HELM): 58 new results, including Gemini 3 Pro, o3, GPT-5… (board)
- WildBench (HELM): 58 new results, including Gemini 3 Pro, o3, GPT-5… (board)
- Omni-MATH (HELM): 58 new results, including Gemini 3 Pro, o3, GPT-5… (board)
- IOI: 20 new results, including Claude Fable 5.1, GPT-6 Astra, Muse Spark 1.3… (board)
- LMArena Text: 2 new results, including GPT-6 Astra, Grok 4.6 (board)
- LMArena Hard Prompts: 2 new results, including GPT-6 Astra, Grok 4.6 (board)
- LMArena Coding: 2 new results, including GPT-6 Astra, Grok 4.6 (board)
- LMArena Agent: 1 new result, including Muse Spark 1.3 (board)
- LMArena WebDev: 1 new result, including DeepSeek V4.1 Flash (board)
Score revisions
- Claude Fable 5.1 (max) on LMArena WebDev: 1764 → 1758. (details)
- Claude Fable 5.1 (max) on LMArena Agent: 14.5 → 13.9. (details)
- Claude Opus 5 on IOI: 91.7% → 84.3%. (details)
- GPT-5.6 Sol (max) on IOI: 86.7% → 91.2%. (details)
- Qwen3 8 (max) on IOI: 73% → 68.9%. (details)
- GLM 5.3 Flash on LMArena Coding: 1534 → 1524. (details)
- Grok 4.6 (high) on LMArena WebDev: 1624 → 1618. (details)
- Kimi K3 (low) on AA Intelligence Index: 34.5 → 30.5. (details)
- Qwen3.8 27B on LMArena Coding: 1494 → 1500. (details)
- GPT-5.6 Luna (max) on IOI: 72.9% → 61.8%. (details)
- Mistral Medium 3.5 on LMArena Agent: -9.3 → -10.1. (details)
- GPT-5.3-Codex (xhigh) on IOI: 43.8% → 53.8%. (details)
- Nemotron 3.5 Lightning 30B A3b Nvfp4 on LMArena Coding: 1432 → 1424. (details)
Speed changes
- Claude Opus 4.7 time to first token changed from 4.55 s to 1.39 s. (details)
- GPT-5.5 time to first answer changed from 6.18 s to 4.01 s. (details)
- Muse Spark 1.1 time to first token changed from 1.51 s to 2.57 s. (details)
- GPT-5.4 Pro output speed changed from 2 tok/s to 1 tok/s. (details)
- GPT-5.4 Pro time to first token changed from 65.61 s to 6.51 s. (details)
- GLM 5.3 Flash time to first token changed from 1.14 s to 1.70 s. (details)
- GPT-5.2-Codex time to first token changed from 3.55 s to 5.93 s. (details)
- DeepSeek V4 Flash time to first token changed from 0.87 s to 1.26 s. (details)
- GLM 4.7 time to first token changed from 0.80 s to 1.71 s. (details)
- DeepSeek V3.1 output speed changed from 18 tok/s to 29 tok/s. (details)
- Inkling Small time to first token changed from 0.86 s to 0.43 s. (details)
- Claude Sonnet 4.5 time to first token changed from 1.10 s to 1.56 s. (details)
- Nemotron 3 Super time to first token changed from 2.94 s to 1.05 s. (details)
- Qwen3.6 27B time to first token changed from 3.25 s to 2.11 s. (details)
- Kimi K2 time to first token changed from 1.52 s to 0.94 s. (details)
- …and 38 more.
Today's top five
- Claude Fable 5.1 — 70.7
- GPT-6 Astra — 70.5
- Claude Fable 5.1 — 70.4
- Claude Fable 5 — 70.2
- Claude Fable 5.1 — 70.1
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.