The week on the leaderboard, 21 Sept 2026 to 27 Sept 2026
GPT-5.2 up 1.6, 12 new models, 23 price cuts, 2211 new benchmark results. What moved on the BenchLeader Index this week and why.
The week of 21 Sept 2026 to 27 Sept 2026, from 7 daily updates.
Three things that happened
- Biggest riser: GPT-5.2, up 1.6 points to 63.5 on the BenchLeader Index as new results landed.
- 12 new models arrived: GLM 5.3 Prime, Qwen3.8 Max Prime, GPT-6 Sol, GPT-6 Luna, GPT-6 Luna Pro and more.
- Busiest benchmark: AA-Omniscience: accuracy with 498 new results; 2211 new results were published in all.
Index movers
- Mistral Small 3.1-3.3
- Muse Spark 1.1-2.9
- Command R+-1.7
- GPT-5.2+1.6
- MiMo-V2.6-Flash+1.3
- Mistral Medium 3.5-1.3
- Nemotron 3 Ultra+1.3
- Mistral Small 3+1.3
| Model | Provider | 21 Sept 2026 | 27 Sept 2026 | Change | |---|---|---:|---:|---:| | Mistral Small 3.1 | Mistral AI | 39.9 | 36.6 | -3.3 | | Muse Spark 1.1 | Meta | 65.2 | 62.3 | -2.9 | | Command R+ | Cohere | 35.7 | 34.0 | -1.7 | | GPT-5.2 | OpenAI | 61.9 | 63.5 | +1.6 | | MiMo-V2.6-Flash | Xiaomi | 59.5 | 60.8 | +1.3 | | Mistral Medium 3.5 | Mistral AI | 51.2 | 49.9 | -1.3 | | Nemotron 3 Ultra | NVIDIA | 48.3 | 49.6 | +1.3 | | Mistral Small 3 | Mistral AI | 38.4 | 39.7 | +1.3 | | Claude Opus 4.8 | Anthropic | 64.1 | 63.0 | -1.1 | | Inkling Small | Thinking Machines | 56.4 | 55.3 | -1.1 |
New models
- GLM 5.3 Prime (Zhipu AI) is now listed at $4.30/M blended; no independent results yet. (details)
- Qwen3.8 Max Prime (Alibaba) is now listed at $6.00/M blended; no independent results yet. (details)
- GPT-6 Sol (no reasoning) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #249. (details)
- GPT-6 Luna (no reasoning) (OpenAI) appeared with 8 benchmark results, entering the BenchLeader Index at #465. (details)
- GPT-6 Luna Pro (OpenAI) is now listed at $0.20/M blended; no independent results yet. (details)
- GPT-6 Sol Pro (OpenAI) is now listed at $4.00/M blended; no independent results yet. (details)
- Claude Opus 5.5 (low) (Anthropic) appeared with 5 benchmark results, entering the BenchLeader Index at #39. (details)
- MiMo-V2.6-Pro (Xiaomi) appeared with 4 benchmark results. (details)
- GPT-6 Astra Pro Max (OpenAI) appeared with 1 benchmark result. (details)
- SWE-2 (max) (Cognition) appeared with 1 benchmark result. (details)
- Grok 4.7 (xhigh) (SpaceXAI) appeared with 6 benchmark results, entering the BenchLeader Index at #72. (details)
- Dots3 Note (max) (Dots Studio) appeared with 1 benchmark result. (details)
Price changes
- Llama 3.1 Nemotron 70B Instruct input price moved from $1.20/M to $0.00/M (100% down).
- Nemotron Nano 12B v2 VL (thinking) input price moved from $0.20/M to $0.00/M (100% down).
- Devstral Small 2 input price moved from $0.10/M to $0.00/M (100% down).
- Ministral 3 14B cost per task moved from 0.1 to 0 (86% down).
- Grok 4.1 (thinking) input price moved from $1.25/M to $0.20/M (84% down).
- Grok 4 Fast (thinking) input price moved from $1.25/M to $0.20/M (84% down).
- Qwen3.5 Flash input price moved from $0.17/M to $0.03/M (83% down).
- Ministral 3 8B cost per task moved from 0.1 to 0 (83% down).
- GLM 5V Turbo output price moved from $22.00/M to $4.00/M (82% down).
- Qwen3 Max (max) output price moved from $7.75/M to $1.43/M (81% down).
- Ministral 3 3B cost per task moved from 0 to 0 (81% down).
- Grok 3 mini (thinking) output price moved from $2.50/M to $0.50/M (80% down).
- Mistral Small 3 output price moved from $0.30/M to $0.08/M (73% down).
- Mistral Small 4 (thinking) cost per task moved from 0 to 0 (67% down).
- Qwen3.5 Plus 2026-04-20 (thinking) output price moved from $1.72/M to $0.69/M (60% down).
Top five as of 11 Oct 2026
- Claude Opus 5.5 — 72.0
- Claude Fable 5.1 — 70.8
- GPT-6 Astra — 70.2
- Claude Opus 5.5 — 70.2
- Gemini 4 Argon — 70.0
Generated from the week's daily updates; every line links to the page where the numbers and their sources can be checked.