Daily digest, 27 Sept 2026
15 price changes, 155 new or updated benchmark results across 22 sources.
BenchLeader refreshed 1521 models from 22 of 22 sources. 15 price changes, 155 new or updated benchmark results across 22 sources.
Price changes
- GLM 5.2 (max) cost per task moved from 1 to 1.5 (53% up). (details)
- Qwen3.7 Plus cost per task moved from 0.3 to 0.2 (26% down). (details)
- Grok 4.3 (high) cost per task moved from 0.2 to 0.2 (24% up). (details)
- Nemotron 3 Ultra 550B A55B (thinking) cost per task moved from 0.5 to 0.6 (9% up). (details)
- GPT-5.4 Mini (xhigh) cost per task moved from 0.4 to 0.4 (9% up). (details)
- Mistral Medium 3.5 cost per task moved from 0.4 to 0.5 (15% up). (details)
- Claude Haiku 4.5 (thinking) cost per task moved from 0.2 to 0.3 (33% up). (details)
- Mercury 2.5 cost per task moved from 0.1 to 0.1 (82% up). (details)
- Grok 4.3 (no reasoning) cost per task moved from 0.1 to 0.2 (24% up). (details)
- Mistral Large 3 cost per task moved from 0.1 to 0 (42% down). (details)
- Mistral Small 3.1 output price moved from $0.17/M to $0.30/M (76% up). (details)
- Mixtral 8x7B input price moved from $0.45/M to $0.70/M (56% up). (details)
- Mistral Medium 3.5 (high) cost per task moved from 0.4 to 0.5 (15% up). (details)
- Mercury 2.5 (high) cost per task moved from 0.1 to 0.1 (82% up). (details)
- Qwen3.7 Plus (no reasoning) cost per task moved from 0.3 to 0.2 (26% down). (details)
New benchmark results
- AudioMC (audio output): 15 new results, including Qwen3 Omni 30B A3B Instruct, Gemini 2.5 Flash Native Audio 12 2025, Gemini 2.5 Flash Native Audio 12 2025… (board)
- DTBench: 6 new results, including Mistral Medium 3.5, Magistral Small 1.2, Mistral Small 3.1… (board)
- Epoch Capabilities Index: 6 new results, including Baichuan1-7B, Falcon 2 11B, Stable Beluga 2… (board)
- EnigmaEval: 5 new results, including GPT-5.6 Sol, Claude Fable 5, Gemini 3.5 Flash… (board)
- SWE-bench Verified (any scaffold): 5 new results, including Claude Sonnet 4.5, GPT-5, Claude Opus 4.5… (board)
- SciCode: 4 new results, including Mistral Medium 3.5, Magistral Small 1.2, Devstral Small 2… (board)
- LiveCodeBench: 4 new results, including Devstral 2, Magistral Small 1.2, Devstral Small 2… (board)
- AA Intelligence Index v4.3.2: 3 new results, including MiMo-V2.6-Flash, Magistral Small 1.2, Mistral Small 4 (board)
- Humanity's Last Exam (AA): 3 new results, including MiMo-V2.6-Flash, Magistral Small 1.2, Mistral Small 4 (board)
- GDPval-AA v2.1: 3 new results, including MiMo-V2.6-Flash, Motif 3, Magistral Small 1.2 (board)
- LMCA: 3 new results, including Mistral Medium 3.5, Command R+, Mistral Small 4 (board)
- MATH Level 5: 3 new results, including Mistral Small 3, Phi 3, Mistral Small 3.1 (board)
Score revisions
- Gemini 3 Pro on SWE-bench Verified (any scaffold): 77.4% → 74.2%. (details)
- Motif 3 on GPQA Diamond (AA): 83.4% → 86.9%. (details)
- Motif 3 on Humanity's Last Exam (AA): 37% → 40.4%. (details)
- Claude Opus 4 on SWE-bench Verified (any scaffold): 73.2% → 67.6%. (details)
- Kimi K2 on SWE-bench Verified (any scaffold): 65.4% → 53.4%. (details)
- Claude Sonnet 4 on SWE-bench Verified (any scaffold): 76.8% → 74.8%. (details)
- Claude 3.7 Sonnet on SWE-bench Verified (any scaffold): 70.4% → 52.8%. (details)
- Qwen3 Coder 480B on SWE-bench Verified (any scaffold): 69.6% → 55.4%. (details)
- GPT-4o on SWE-bench Verified (any scaffold): 38.8% → 21.6%. (details)
- Gemini 2.0 Flash on SWE-bench Verified (any scaffold): 52.2% → 13.5%. (details)
- Qwen2.5 Coder 32B Instruct on SWE-bench Verified (any scaffold): 47% → 38%. (details)
- Mistral Small 3.1 on GPQA Diamond: 41.9% → 47.5%. (details)
- Mistral Small 3.1 on OTIS Mock AIME: 3.9% → 5.8%. (details)
- Mixtral 8x7B on MATH Level 5: 9.3% → 10%. (details)
- Baichuan2-13B on MMLU: 55.1% → 59.2%. (details)
- Baichuan2-13B on GSM8K: 45.7% → 52.8%. (details)
- Baichuan2-13B on BIG-Bench Hard: 47.2% → 49%. (details)
Speed changes
- Claude Opus 5.5 output speed changed from 61 tok/s to 95 tok/s. (details)
- Claude Fable 5 time to first answer changed from 153.30 s to 96.96 s. (details)
- Claude Opus 5 time to first token changed from 2.87 s to 4.03 s. (details)
- GPT-5.5 time to first answer changed from 65.38 s to 35.33 s. (details)
- GPT-5.5 response time changed from 70.03 s to 39.95 s. (details)
- Claude Opus 4.7 time to first token changed from 3.48 s to 1.38 s. (details)
- GPT-5.2 time to first token changed from 2.33 s to 3.26 s. (details)
- DeepSeek V4 Pro time to first token changed from 1.52 s to 2.32 s. (details)
- GLM 5.2 time to first token changed from 0.87 s to 1.21 s. (details)
- GPT-5.4 Pro output speed changed from 6 tok/s to 1 tok/s. (details)
- GPT-5.4 Pro time to first token changed from 62.03 s to 5.67 s. (details)
- GLM 5.3 Flash time to first token changed from 1.05 s to 1.54 s. (details)
- Claude Opus 4.8 time to first answer changed from 15.30 s to 9.41 s. (details)
- Grok 4.7 output speed changed from 57 tok/s to 82 tok/s. (details)
- Grok 4.7 time to first answer changed from 2.81 s to 48.09 s. (details)
- …and 73 more.
Today's top five
- Claude Opus 5.5 — 72.0
- Claude Fable 5.1 — 70.8
- GPT-6 Astra — 70.2
- Claude Opus 5.5 — 70.2
- Gemini 4 Argon — 70.0
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.