BenchLeader

Daily digest, 24 Sept 2026

2 new models, 12 price changes, 21 new or updated benchmark results across 22 sources.

BenchLeader refreshed 1550 models from 22 of 22 sources. 2 new models, 12 price changes, 21 new or updated benchmark results across 22 sources.

New models

  • GLM 5.3 Prime (Zhipu AI) is now listed at $4.30/M blended; no independent results yet. (details)
  • Qwen3.8 Max Prime (Alibaba) is now listed at $6.00/M blended; no independent results yet. (details)

Price changes

  • Trinity Large (thinking) output price moved from $0.90/M to $0.80/M (11% down). (details)
  • Llama 3.1 Nemotron 70B Instruct input price moved from $1.20/M to $0.00/M (100% down). (details)
  • Llama 3.1 Nemotron 70B Instruct output price moved from $1.20/M to $0.00/M (100% down). (details)
  • Nemotron Nano 12B v2 VL (thinking) input price moved from $0.20/M to $0.00/M (100% down). (details)
  • Nemotron Nano 12B v2 VL (thinking) output price moved from $0.60/M to $0.00/M (100% down). (details)
  • Mistral Small 3.1 input price moved from $0.10/M to $0.11/M (10% up). (details)
  • Mistral Small 3.1 output price moved from $0.30/M to $0.17/M (43% down). (details)
  • Mistral Small 3.1 cost per task moved from 0 to 0 (7% up). (details)
  • Mistral Small 3 input price moved from $0.10/M to $0.05/M (50% down). (details)
  • Mistral Small 3 output price moved from $0.30/M to $0.08/M (73% down). (details)
  • Mistral 7B v0.3 input price moved from $0.25/M to $0.15/M (40% down). (details)
  • Mistral 7B v0.3 output price moved from $0.25/M to $0.20/M (20% down). (details)

New benchmark results

  • LMArena WebDev: 2 new results, including GPT-6 Sol, Claude Opus 5.5 (board)
  • CyberBench: 2 new results, including MiMo-V2.6-Pro, MiMo-V2.6-Flash (board)
  • MysteryMechanism: 2 new results, including MiMo-V2.6-Pro, MiMo-V2.6-Flash (board)
  • Public Benefits Bench: 2 new results, including MiMo-V2.6-Pro, MiMo-V2.6-Flash (board)
  • Tax Agent Bench: 2 new results, including MiMo-V2.6-Pro, MiMo-V2.6-Flash (board)
  • Terminal-Bench 4.0 (Vals): 2 new results, including MiMo-V2.6-Pro, MiMo-V2.6-Flash (board)
  • EBR-bench: 1 new result, including GPT-6 Sol (board)
  • IOI: 1 new result, including MiMo-V2.6-Pro (board)
  • ProgramBench: 1 new result, including MiMo-V2.6-Pro (board)
  • Terminal-Bench 4.0 (AA): 1 new result, including Apodex 1.1 (board)
  • Terminal-Bench Science: 1 new result, including MiMo-V2.6-Flash (board)

Score revisions

  • Kimi K3 (max) on AA-Briefcase: 1510 → 1504. (details)
  • GLM 5.3 (max) on AA-Briefcase: 1525 → 1516. (details)
  • Grok 4.6 (high) on LMArena WebDev: 1616 → 1624. (details)
  • MiMo-V2.6-Pro on Finance Agent v2: 58.3% → 57.3%. (details)

Speed changes

  • GPT-6 Sol time to first token changed from 2.35 s to 4.65 s. (details)
  • GPT-5.1-Codex output speed changed from 24 tok/s to 35 tok/s. (details)
  • Ling 3.0 Flash VL time to first token changed from 0.90 s to 1.41 s. (details)
  • Kimi K2 Thinking time to first token changed from 1.84 s to 0.99 s. (details)
  • Nemotron 3 Nano 30B A3B output speed changed from 129 tok/s to 184 tok/s. (details)
  • Mistral Medium 3 output speed changed from 153 tok/s to 65 tok/s. (details)
  • Mistral Small 3.2 output speed changed from 157 tok/s to 24 tok/s. (details)
  • Nova Pro output speed changed from 30 tok/s to 12 tok/s. (details)
  • Qwen3.5-9B output speed changed from 26 tok/s to 37 tok/s. (details)
  • Nemotron Nano 12B v2 VL time to first answer changed from 2.92 s to 1.10 s. (details)
  • Nemotron Nano 12B v2 VL time to first token changed from 2.92 s to 1.10 s. (details)
  • Nemotron Nano 12B v2 VL response time changed from 5.60 s to 3.56 s. (details)
  • Mistral Small 4 output speed changed from 154 tok/s to 93 tok/s. (details)
  • Hy-MT2-1.8B output speed changed from 98 tok/s to 38 tok/s. (details)
  • Hy-MT2-1.8B time to first token changed from 0.34 s to 0.59 s. (details)
  • …and 5 more.

Today's top five

  1. GPT-6 Astra — 71.4
  2. GPT-6 Astra — 71.1
  3. Claude Fable 5.1 — 70.9
  4. Claude Fable 5.1 — 70.8
  5. Claude Fable 5.1 — 70.8

This digest is generated automatically from the day's data changes. Every line links to the page where you can check the numbers and their source.

Get this in your inbox

No ads, no tracking, unsubscribe in one click.
What to receive