Daily digest, 17 Sept 2026
3 new models, 17 price changes, 8 new or updated benchmark results across 20 sources.
BenchLeader refreshed 1271 models from 20 of 20 sources. 3 new models, 17 price changes, 8 new or updated benchmark results across 20 sources.
New models
- Union Alpha (Unknown) appeared with 1 benchmark result. (details)
- Jamba Large (AI21 Labs) is now listed at $3.50/M blended; no independent results yet. (details)
- Jamba Mini (AI21 Labs) is now listed at $0.25/M blended; no independent results yet. (details)
Price changes
- Qwen3.5 Plus 2026-04-20 (thinking) input price moved from $0.57/M to $0.29/M (50% down). (details)
- Qwen3.5 Plus 2026-04-20 (thinking) output price moved from $3.44/M to $1.72/M (50% down). (details)
- Qwen3 Max (max) input price moved from $0.86/M to $1.29/M (50% up). (details)
- Qwen3 Max (max) output price moved from $3.44/M to $7.75/M (125% up). (details)
- Qwen3.5 Flash output price moved from $1.72/M to $1.03/M (40% down). (details)
- Llama 3.3 70B input price moved from $0.66/M to $0.71/M (8% up). (details)
- Qwen3.5 Plus 2026-04-20 input price moved from $0.57/M to $0.29/M (50% down). (details)
- Qwen3.5 Plus 2026-04-20 output price moved from $3.44/M to $1.72/M (50% down). (details)
- Gemini Pro input price moved from $1.25/M to $2.00/M (60% up). (details)
- Gemini Pro output price moved from $10.00/M to $12.00/M (20% up). (details)
- Qwen3.5 Flash (no reasoning) output price moved from $1.72/M to $1.03/M (40% down). (details)
- Muse Glimmer 30B input price moved from $0.35/M to $0.30/M (14% down). (details)
- Muse Glimmer 30B output price moved from $1.50/M to $1.10/M (27% down). (details)
- Qwen-VL OCR input price moved from $0.72/M to $0.04/M (94% down). (details)
- Qwen-VL OCR output price moved from $0.72/M to $0.07/M (90% down). (details)
- Qwen3 Coder Plus input price moved from $1.00/M to $0.57/M (43% down). (details)
- Qwen3 Coder Plus output price moved from $5.00/M to $2.30/M (54% down). (details)
New benchmark results
- PRBench Finance: 1 new result, including Qwen3 235B A22B 2507 (board)
- PRBench Legal: 1 new result, including Qwen3 235B A22B 2507 (board)
- SWE-Bench Pro: 1 new result, including Llama 4 Maverick (board)
- AA Intelligence Index: 1 new result, including Ling 3.0 Flash Fin (board)
- Humanity's Last Exam (AA): 1 new result, including Ling 3.0 Flash Fin (board)
- AA-LCR: 1 new result, including Ling 3.0 Flash Fin (board)
- SciCode (AA): 1 new result, including Ling 3.0 Flash Fin (board)
- AA-Omniscience: 1 new result, including Ling 3.0 Flash Fin (board)
Speed changes
- Claude Fable 5 time to first answer changed from 96.09 s to 129.84 s. (details)
- Claude Fable 5.1 output speed changed from 47 tok/s to 68 tok/s. (details)
- GPT-6 Astra output speed changed from 31 tok/s to 53 tok/s. (details)
- Claude Fable 5 output speed changed from 49 tok/s to 68 tok/s. (details)
- GPT-5.6 Sol output speed changed from 43 tok/s to 73 tok/s. (details)
- GPT-5.5 output speed changed from 49 tok/s to 91 tok/s. (details)
- Gemini 3.8 Flash output speed changed from 74 tok/s to 345 tok/s. (details)
- Gemini 3.8 Flash time to first answer changed from 14.66 s to 30.49 s. (details)
- Muse Spark 1.3 output speed changed from 95 tok/s to 255 tok/s. (details)
- GPT-5.4 output speed changed from 53 tok/s to 136 tok/s. (details)
- Gemini 3.7 Flash output speed changed from 85 tok/s to 328 tok/s. (details)
- Muse Spark 1.2 output speed changed from 112 tok/s to 208 tok/s. (details)
- Gemini 3.1 Pro output speed changed from 90 tok/s to 122 tok/s. (details)
- Gemini 3.5 Flash output speed changed from 116 tok/s to 227 tok/s. (details)
- GPT-5.1 time to first answer changed from 28.90 s to 43.42 s. (details)
- …and 92 more.
Today's top five
- Claude Fable 5.1 — 70.5
- Claude Fable 5.1 — 70.4
- Claude Fable 5 — 70.1
- GPT-6 Astra — 70.1
- Claude Fable 5.1 — 69.9
This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.