BenchLeader

The week on the leaderboard, 21 Sept 2026 to 27 Sept 2026

GPT-5.2 up 1.6, 12 new models, 23 price cuts, 2211 new benchmark results. What moved on the BenchLeader Index this week and why.

The week of 21 Sept 2026 to 27 Sept 2026, from 7 daily updates.

Three things that happened

Index movers

Change in BenchLeader Index, 2026-09-21 to 2026-09-27; best configuration per model.

| Model | Provider | 21 Sept 2026 | 27 Sept 2026 | Change | |---|---|---:|---:|---:| | Mistral Small 3.1 | Mistral AI | 39.9 | 36.6 | -3.3 | | Muse Spark 1.1 | Meta | 65.2 | 62.3 | -2.9 | | Command R+ | Cohere | 35.7 | 34.0 | -1.7 | | GPT-5.2 | OpenAI | 61.9 | 63.5 | +1.6 | | MiMo-V2.6-Flash | Xiaomi | 59.5 | 60.8 | +1.3 | | Mistral Medium 3.5 | Mistral AI | 51.2 | 49.9 | -1.3 | | Nemotron 3 Ultra | NVIDIA | 48.3 | 49.6 | +1.3 | | Mistral Small 3 | Mistral AI | 38.4 | 39.7 | +1.3 | | Claude Opus 4.8 | Anthropic | 64.1 | 63.0 | -1.1 | | Inkling Small | Thinking Machines | 56.4 | 55.3 | -1.1 |

New models

  • GLM 5.3 Prime (Zhipu AI) is now listed at $4.30/M blended; no independent results yet. (details)
  • Qwen3.8 Max Prime (Alibaba) is now listed at $6.00/M blended; no independent results yet. (details)
  • GPT-6 Sol (no reasoning) (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #249. (details)
  • GPT-6 Luna (no reasoning) (OpenAI) appeared with 8 benchmark results, entering the BenchLeader Index at #465. (details)
  • GPT-6 Luna Pro (OpenAI) is now listed at $0.20/M blended; no independent results yet. (details)
  • GPT-6 Sol Pro (OpenAI) is now listed at $4.00/M blended; no independent results yet. (details)
  • Claude Opus 5.5 (low) (Anthropic) appeared with 5 benchmark results, entering the BenchLeader Index at #39. (details)
  • MiMo-V2.6-Pro (Xiaomi) appeared with 4 benchmark results. (details)
  • GPT-6 Astra Pro Max (OpenAI) appeared with 1 benchmark result. (details)
  • SWE-2 (max) (Cognition) appeared with 1 benchmark result. (details)
  • Grok 4.7 (xhigh) (SpaceXAI) appeared with 6 benchmark results, entering the BenchLeader Index at #72. (details)
  • Dots3 Note (max) (Dots Studio) appeared with 1 benchmark result. (details)

Price changes

  • Llama 3.1 Nemotron 70B Instruct input price moved from $1.20/M to $0.00/M (100% down).
  • Nemotron Nano 12B v2 VL (thinking) input price moved from $0.20/M to $0.00/M (100% down).
  • Devstral Small 2 input price moved from $0.10/M to $0.00/M (100% down).
  • Ministral 3 14B cost per task moved from 0.1 to 0 (86% down).
  • Grok 4.1 (thinking) input price moved from $1.25/M to $0.20/M (84% down).
  • Grok 4 Fast (thinking) input price moved from $1.25/M to $0.20/M (84% down).
  • Qwen3.5 Flash input price moved from $0.17/M to $0.03/M (83% down).
  • Ministral 3 8B cost per task moved from 0.1 to 0 (83% down).
  • GLM 5V Turbo output price moved from $22.00/M to $4.00/M (82% down).
  • Qwen3 Max (max) output price moved from $7.75/M to $1.43/M (81% down).
  • Ministral 3 3B cost per task moved from 0 to 0 (81% down).
  • Grok 3 mini (thinking) output price moved from $2.50/M to $0.50/M (80% down).
  • Mistral Small 3 output price moved from $0.30/M to $0.08/M (73% down).
  • Mistral Small 4 (thinking) cost per task moved from 0 to 0 (67% down).
  • Qwen3.5 Plus 2026-04-20 (thinking) output price moved from $1.72/M to $0.69/M (60% down).

Top five as of 11 Oct 2026

  1. Claude Opus 5.5 — 72.0
  2. Claude Fable 5.1 — 70.8
  3. GPT-6 Astra — 70.2
  4. Claude Opus 5.5 — 70.2
  5. Gemini 4 Argon — 70.0

Generated from the week's daily updates; every line links to the page where the numbers and their sources can be checked.

Get the weekly review by email

No ads, no tracking, unsubscribe in one click.
What to receive