The week on the leaderboard, 7 Sept 2026 to 13 Sept 2026
Gemma 4 31B up 7.4, 408 new models, 159 price cuts, 4986 new benchmark results. What moved on the BenchLeader Index this week and why.
The week of 7 Sept 2026 to 13 Sept 2026 in 6779 recorded changes across 6 daily updates.
Three things that happened
- Biggest riser: Gemma 4 31B, up 7.4 points to 53.8 on the BenchLeader Index as new results landed.
- 408 new models arrived: AU Anthropic Claude Opus 4.6, AU Anthropic Claude Sonnet 4.6, Claude Fable 5 (EU), Claude Fable 5 (Global), Claude Fable 5 (US) and more.
- Busiest benchmark: aa_intelligence_index with 377 new results; 4986 new results were published in all.
Index movers
- Qwen3 235B A22B-10.8
- Qwen3 30B A3B-10.4
- Gemma 4 31B+7.4
- Mistral Large 3-7.1
- Qwen3 32B-6.9
- MiniMax M3+6.7
- Qwen3-6.7
- Mistral Medium 3.5-6.7
| Model | Provider | 7 Sept 2026 | 13 Sept 2026 | Change | |---|---|---:|---:|---:| | Qwen3 235B A22B | Alibaba | 53.8 | 43.0 | -10.8 | | Qwen3 30B A3B | Alibaba | 50.8 | 40.4 | -10.4 | | Gemma 4 31B | Google | 46.4 | 53.8 | +7.4 | | Mistral Large 3 | Mistral AI | 52.1 | 45.0 | -7.1 | | Qwen3 32B | Alibaba | 51.4 | 44.5 | -6.9 | | MiniMax M3 | MiniMax | 49.3 | 56.0 | +6.7 | | Qwen3 | Alibaba | 55.0 | 48.3 | -6.7 | | Mistral Medium 3.5 | Mistral AI | 48.3 | 41.6 | -6.7 | | GPT-5 mini | OpenAI | 50.5 | 56.6 | +6.1 | | MiniMax M2.7 | MiniMax | 49.3 | 55.4 | +6.1 |
New models
- AU Anthropic Claude Opus 4.6 (Amazon) is now listed at $33.00/M blended; no independent results yet. (details)
- AU Anthropic Claude Sonnet 4.6 (Amazon) is now listed at $6.60/M blended; no independent results yet. (details)
- Claude Fable 5 (EU) (Amazon) is now listed at $22.00/M blended; no independent results yet. (details)
- Claude Fable 5 (Global) (Amazon) is now listed at $20.00/M blended; no independent results yet. (details)
- Claude Fable 5 (US) (Amazon) is now listed at $20.00/M blended; no independent results yet. (details)
- Claude Fable 5.1 (Global) (Amazon) is now listed at $20.00/M blended; no independent results yet. (details)
- Claude Fable 5.1 (US) (Amazon) is now listed at $22.00/M blended; no independent results yet. (details)
- Claude Haiku 4.5 (AU) (Amazon) is now listed at $2.00/M blended; no independent results yet. (details)
- Claude Haiku 4.5 (EU) (Amazon) is now listed at $2.20/M blended; no independent results yet. (details)
- Claude Haiku 4.5 (Global) (Amazon) is now listed at $2.00/M blended; no independent results yet. (details)
- Claude Haiku 4.5 (JP) (Amazon) is now listed at $2.00/M blended; no independent results yet. (details)
- Claude Haiku 4.5 (US) (Amazon) is now listed at $2.00/M blended; no independent results yet. (details)
- Claude Opus 4.1 (US) (Amazon) is now listed at $30.00/M blended; no independent results yet. (details)
- Claude Opus 4.5 (EU) (Amazon) is now listed at $11.00/M blended; no independent results yet. (details)
- Claude Opus 4.5 (Global) (Amazon) is now listed at $10.00/M blended; no independent results yet. (details)
- Claude Opus 4.5 (US) (Amazon) is now listed at $10.00/M blended; no independent results yet. (details)
- Claude Opus 4.6 (EU) (Amazon) is now listed at $11.00/M blended; no independent results yet. (details)
- Claude Opus 4.6 (Global) (Amazon) is now listed at $10.00/M blended; no independent results yet. (details)
- Claude Opus 4.6 (US) (Amazon) is now listed at $10.00/M blended; no independent results yet. (details)
- Claude Opus 4.7 (AU) (Amazon) is now listed at $11.00/M blended; no independent results yet. (details)
- Claude Opus 4.7 (EU) (Amazon) is now listed at $11.00/M blended; no independent results yet. (details)
- Claude Opus 4.7 (Global) (Amazon) is now listed at $10.00/M blended; no independent results yet. (details)
- Claude Opus 4.7 (JP) (Amazon) is now listed at $10.00/M blended; no independent results yet. (details)
- Claude Opus 4.7 (US) (Amazon) is now listed at $10.00/M blended; no independent results yet. (details)
- Claude Opus 4.8 (AU) (Amazon) is now listed at $10.00/M blended; no independent results yet. (details)
- Claude Opus 4.8 (EU) (Amazon) is now listed at $11.00/M blended; no independent results yet. (details)
- Claude Opus 4.8 (Global) (Amazon) is now listed at $10.00/M blended; no independent results yet. (details)
- Claude Opus 4.8 (JP) (Amazon) is now listed at $10.00/M blended; no independent results yet. (details)
- Claude Opus 4.8 (US) (Amazon) is now listed at $10.00/M blended; no independent results yet. (details)
- Claude Opus 5 (AU) (Amazon) is now listed at $10.00/M blended; no independent results yet. (details)
- Claude Opus 5 (EU) (Amazon) is now listed at $11.00/M blended; no independent results yet. (details)
- Claude Opus 5 (Global) (Amazon) is now listed at $10.00/M blended; no independent results yet. (details)
- Claude Opus 5 (JP) (Amazon) is now listed at $10.00/M blended; no independent results yet. (details)
- Claude Opus 5 (US) (Amazon) is now listed at $10.00/M blended; no independent results yet. (details)
- Claude Sonnet 4.5 (AU) (Amazon) is now listed at $6.00/M blended; no independent results yet. (details)
- Claude Sonnet 4.5 (EU) (Amazon) is now listed at $6.60/M blended; no independent results yet. (details)
- Claude Sonnet 4.5 (Global) (Amazon) is now listed at $6.00/M blended; no independent results yet. (details)
- Claude Sonnet 4.5 (JP) (Amazon) is now listed at $6.00/M blended; no independent results yet. (details)
- Claude Sonnet 4.5 (US) (Amazon) is now listed at $6.00/M blended; no independent results yet. (details)
- Claude Sonnet 4.6 (EU) (Amazon) is now listed at $6.60/M blended; no independent results yet. (details)
- Claude Sonnet 4.6 (Global) (Amazon) is now listed at $6.00/M blended; no independent results yet. (details)
- Claude Sonnet 4.6 (JP) (Amazon) is now listed at $6.00/M blended; no independent results yet. (details)
- Claude Sonnet 4.6 (US) (Amazon) is now listed at $6.00/M blended; no independent results yet. (details)
- Claude Sonnet 5 (AU) (Amazon) is now listed at $4.00/M blended; no independent results yet. (details)
- Claude Sonnet 5 (EU) (Amazon) is now listed at $4.40/M blended; no independent results yet. (details)
- Claude Sonnet 5 (Global) (Amazon) is now listed at $4.00/M blended; no independent results yet. (details)
- Claude Sonnet 5 (JP) (Amazon) is now listed at $4.00/M blended; no independent results yet. (details)
- Claude Sonnet 5 (US) (Amazon) is now listed at $4.00/M blended; no independent results yet. (details)
- DeepSeek-R1 (US) (Amazon) is now listed at $2.36/M blended; no independent results yet. (details)
- GLM-4.7 (Amazon) is now listed at $1.00/M blended; no independent results yet. (details)
- GLM-4.7-Flash (Amazon) is now listed at $0.15/M blended; no independent results yet. (details)
- GLM-5 (Amazon) is now listed at $1.55/M blended; no independent results yet. (details)
- GPT-5.6 Luna (Global) (Amazon) is now listed at $0.45/M blended; no independent results yet. (details)
- GPT-5.6 Luna (India) (Amazon) is now listed at $0.49/M blended; no independent results yet. (details)
- GPT-5.6 Luna (US) (Amazon) is now listed at $0.49/M blended; no independent results yet. (details)
- GPT-5.6 Sol (Global) (Amazon) is now listed at $8.00/M blended; no independent results yet. (details)
- GPT-5.6 Sol (US) (Amazon) is now listed at $8.80/M blended; no independent results yet. (details)
- GPT-5.6 Terra (Global) (Amazon) is now listed at $4.50/M blended; no independent results yet. (details)
- GPT-5.6 Terra (India) (Amazon) is now listed at $4.95/M blended; no independent results yet. (details)
- GPT-5.6 Terra (US) (Amazon) is now listed at $4.95/M blended; no independent results yet. (details)
- GPT-6 Astra (Global) (Amazon) is now listed at $20.00/M blended; no independent results yet. (details)
- GPT-6 Astra (US) (Amazon) is now listed at $22.00/M blended; no independent results yet. (details)
- gpt-oss-120b (GovCloud) (Amazon) is now listed at $0.32/M blended; no independent results yet. (details)
- gpt-oss-20b (GovCloud) (Amazon) is now listed at $0.15/M blended; no independent results yet. (details)
- Grok 4.6 (Global) (Amazon) is now listed at $3.00/M blended; no independent results yet. (details)
- Grok 4.6 (US) (Amazon) is now listed at $3.30/M blended; no independent results yet. (details)
- Llama 3.1 70B Instruct (US) (Amazon) is now listed at $0.72/M blended; no independent results yet. (details)
- Llama 3.1 8B Instruct (US) (Amazon) is now listed at $0.22/M blended; no independent results yet. (details)
- Llama 3.3 70B Instruct (US) (Amazon) is now listed at $0.72/M blended; no independent results yet. (details)
- Llama 4 Maverick 17B Instruct (US) (Amazon) is now listed at $0.42/M blended; no independent results yet. (details)
- Llama 4 Scout 17B Instruct (US) (Amazon) is now listed at $0.29/M blended; no independent results yet. (details)
- mistral-nemotron (NVIDIA) is now listed; no independent results yet. (details)
- Nova 2 Lite (EU) (Amazon) is now listed at $1.07/M blended; no independent results yet. (details)
- Nova 2 Lite (Global) (Amazon) is now listed at $0.85/M blended; no independent results yet. (details)
- Nova 2 Lite (JP) (Amazon) is now listed at $1.13/M blended; no independent results yet. (details)
- Nova 2 Lite (US) (Amazon) is now listed at $0.94/M blended; no independent results yet. (details)
- Nova Lite (APAC) (Amazon) is now listed at $0.11/M blended; no independent results yet. (details)
- Nova Lite (CA) (Amazon) is now listed at $0.11/M blended; no independent results yet. (details)
- Nova Lite (EU) (Amazon) is now listed at $0.12/M blended; no independent results yet. (details)
- Nova Lite (US) (Amazon) is now listed at $0.10/M blended; no independent results yet. (details)
- Nova Micro (APAC) (Amazon) is now listed at $0.07/M blended; no independent results yet. (details)
- Nova Micro (EU) (Amazon) is now listed at $0.07/M blended; no independent results yet. (details)
- Nova Micro (US) (Amazon) is now listed at $0.06/M blended; no independent results yet. (details)
- Nova Premier (US) (Amazon) is now listed at $5.00/M blended; no independent results yet. (details)
- Nova Pro (APAC) (Amazon) is now listed at $1.47/M blended; no independent results yet. (details)
- Nova Pro (EU) (Amazon) is now listed at $1.61/M blended; no independent results yet. (details)
- Nova Pro (US) (Amazon) is now listed at $1.40/M blended; no independent results yet. (details)
- Palmyra X4 (Amazon) is now listed at $4.38/M blended; no independent results yet. (details)
- Palmyra X4 (US) (Amazon) is now listed at $4.38/M blended; no independent results yet. (details)
- Palmyra X5 (Amazon) is now listed at $1.95/M blended; no independent results yet. (details)
- Palmyra X5 (US) (Amazon) is now listed at $1.95/M blended; no independent results yet. (details)
- Pixtral Large (25.02) (EU) (Amazon) is now listed at $3.00/M blended; no independent results yet. (details)
- Pixtral Large (25.02) (US) (Amazon) is now listed at $3.00/M blended; no independent results yet. (details)
- Granite 3.3 8B (IBM) appeared with 3 benchmark results. (details)
- Granite 4.0 Small (IBM) appeared with 3 benchmark results. (details)
- Marin 8B (Unknown) appeared with 3 benchmark results. (details)
- Mistral v0 3 7B (Mistral AI) appeared with 3 benchmark results. (details)
- Olmo 2 13B November 2024 (Ai2) appeared with 3 benchmark results. (details)
- Olmo 2 32B March 2025 (Ai2) appeared with 3 benchmark results. (details)
- Olmo 2 7B November 2024 (Ai2) appeared with 3 benchmark results. (details)
- Olmoe 1B 7B January 2025 (Ai2) appeared with 3 benchmark results. (details)
- Palmyra Fin (Writer) appeared with 3 benchmark results. (details)
- Palmyra Med (Writer) appeared with 3 benchmark results. (details)
- Agnes 3.0 Flash (Sapiens AI) appeared with 3 benchmark results. (details)
- MiMo-V2-Omni-0327 (Xiaomi) appeared with 6 benchmark results, entering the BenchLeader Index at #135. (details)
- Qwen3 (max) (Alibaba) appeared with 5 benchmark results, entering the BenchLeader Index at #192. (details)
- DeepSeek V4 Flash 0731 (DeepSeek) appeared with 5 benchmark results, entering the BenchLeader Index at #277. (details)
- Kimi K2 0905 (Moonshot AI) appeared with 11 benchmark results, entering the BenchLeader Index at #280. (details)
- Qwen3 235B A22B 2507 (Alibaba) appeared with 10 benchmark results, entering the BenchLeader Index at #281. (details)
- Gemini 2.5 Flash 09 2025 (Google) appeared with 20 benchmark results, entering the BenchLeader Index at #283. (details)
- DeepSeek R1 0528 (DeepSeek) appeared with 13 benchmark results, entering the BenchLeader Index at #313. (details)
- Qwen3 30B A3B 2507 (Alibaba) appeared with 11 benchmark results, entering the BenchLeader Index at #377. (details)
- GPT 4 0125 (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #396. (details)
- Magistral Medium 1.2 (Mistral AI) appeared with 16 benchmark results, entering the BenchLeader Index at #403. (details)
- GPT 4 0314 (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #485. (details)
- qwen3-4b-instruct-2507 (Alibaba) appeared with 8 benchmark results, entering the BenchLeader Index at #501. (details)
- Magistral Small 1.2 (Mistral AI) appeared with 17 benchmark results, entering the BenchLeader Index at #505. (details)
- Mistral Large (Mistral AI) appeared with 6 benchmark results, entering the BenchLeader Index at #514. (details)
- GPT 3.5 Turbo 1106 (OpenAI) appeared with 5 benchmark results, entering the BenchLeader Index at #545. (details)
- GPT 4 0613 (OpenAI) appeared with 7 benchmark results, entering the BenchLeader Index at #549. (details)
- GPT 3.5 Turbo 0125 (OpenAI) appeared with 8 benchmark results, entering the BenchLeader Index at #607. (details)
- Mistral Small 3.1 (Mistral AI) appeared with 12 benchmark results, entering the BenchLeader Index at #616. (details)
- Mistral Small 2402 (Mistral AI) appeared with 7 benchmark results, entering the BenchLeader Index at #623. (details)
- DeepSeek V4.1 Flash (DeepSeek) appeared with 4 benchmark results. (details)
- Ling-3.0-flash-VL (Ant Group) appeared with 4 benchmark results. (details)
- Mistral Small 3.1 24B 2503 (Mistral AI) appeared with 4 benchmark results. (details)
- Athene 70B 0725 (Nexusflow) appeared with 3 benchmark results. (details)
- Command R (08-2024) (Cohere) appeared with 3 benchmark results. (details)
- Command R+ (08-2024) (Cohere) appeared with 3 benchmark results. (details)
- Deepseek v2 5 1210 (DeepSeek) appeared with 3 benchmark results. (details)
- Devstral 2 2512 (Mistral AI) appeared with 3 benchmark results. (details)
- Ernie 5.0 0110 (Baidu) appeared with 3 benchmark results. (details)
- Ernie 5.0 1022 (Baidu) appeared with 3 benchmark results. (details)
- Ernie 5.0 1203 (Baidu) appeared with 3 benchmark results. (details)
- Gemini Advanced 0514 (Google) appeared with 3 benchmark results. (details)
- Glm 4 0520 (Zhipu AI) appeared with 3 benchmark results. (details)
- Glm 4 Plus 0111 (Zhipu AI) appeared with 3 benchmark results. (details)
- Hunyuan Turbo 0110 (Tencent) appeared with 3 benchmark results. (details)
- K-EXAONE 2.0 (LG AI Research) appeared with 3 benchmark results. (details)
- Magistral Medium (Mistral AI) appeared with 3 benchmark results. (details)
- Mistral Small 3 (Mistral AI) appeared with 3 benchmark results. (details)
- Mistral Small 3 (Mistral AI) appeared with 3 benchmark results. (details)
- Motif 3 (Beta) (Motif Technologies) appeared with 3 benchmark results. (details)
- Openchat 3.5 0106 (OpenChat) appeared with 3 benchmark results. (details)
- Phi 3 (medium) (Microsoft) appeared with 3 benchmark results. (details)
- Phi 3 Mini June 2024 (Microsoft) appeared with 3 benchmark results. (details)
- Qwen Max 0919 (Alibaba) appeared with 3 benchmark results. (details)
- Qwen2 5 Plus 1127 (Alibaba) appeared with 3 benchmark results. (details)
- Qwen3 5 (max) (Alibaba) appeared with 3 benchmark results. (details)
- DeepSeek V4 Pro 0813 (DeepSeek) appeared with 2 benchmark results. (details)
- Devstral Small 2 (Mistral AI) appeared with 2 benchmark results. (details)
- Gemini 1206 (Google) appeared with 2 benchmark results. (details)
- Magistral Small 1.0 (Mistral AI) appeared with 2 benchmark results. (details)
- phi-3-medium 14B (medium) (Microsoft) appeared with 2 benchmark results. (details)
- Qwen Qwen3 235B A22b 2507 (Alibaba) appeared with 2 benchmark results. (details)
- Chatgpt 4o 01.29 2025 (OpenAI) appeared with 1 benchmark result. (details)
- Codestral 2405 (Mistral AI) appeared with 1 benchmark result. (details)
- Codestral 2508 (Mistral AI) appeared with 1 benchmark result. (details)
- Command A Plus (Cohere) appeared with 1 benchmark result. (details)
- Command R+ (Cohere) appeared with 1 benchmark result. (details)
- Devstral Medium (Mistral AI) appeared with 1 benchmark result. (details)
- Ernie 5.0 1220 (Baidu) appeared with 1 benchmark result. (details)
- Molmo 72B 0924 (Ai2) appeared with 1 benchmark result. (details)
- Molmo 7B D 0924 (Ai2) appeared with 1 benchmark result. (details)
- Qwen Vl Max 1119 (Alibaba) appeared with 1 benchmark result. (details)
- Qwen3 235B 2507 (thinking) (Alibaba) appeared with 1 benchmark result. (details)
- Qwen3.8 Max (0902) (Alibaba) appeared with 1 benchmark result. (details)
- R1 1776 (Perplexity) appeared with 1 benchmark result. (details)
- Command A 2025 (thinking) (Cohere) is now listed at $4.38/M blended; no independent results yet. (details)
- Command A Translate (Cohere) is now listed at $4.38/M blended; no independent results yet. (details)
- Command A Vision (Cohere) is now listed at $4.38/M blended; no independent results yet. (details)
- Command R7B (12-2024) (Cohere) is now listed at $0.07/M blended; no independent results yet. (details)
- Command R7B Arabic (Cohere) is now listed at $0.07/M blended; no independent results yet. (details)
- Deep Research 04 2026 (Google) is now listed at $4.50/M blended; no independent results yet. (details)
- Deep Research Max 04 2026 (Google) is now listed at $4.50/M blended; no independent results yet. (details)
- Devstral Small (Mistral AI) is now listed at $0.15/M blended; no independent results yet. (details)
- Frogboss 32B 2510 (Unknown) is now listed; no independent results yet. (details)
- Frogmini 14B 2510 (Unknown) is now listed; no independent results yet. (details)
- Gemini 2.5 Computer Use 10 2025 (Google) is now listed at $3.44/M blended; no independent results yet. (details)
- GPT-3.5 Turbo (older v0613) (OpenAI) is now listed at $1.25/M blended; no independent results yet. (details)
- Ministral 3 3B 2512 (Mistral AI) is now listed at $0.10/M blended; no independent results yet. (details)
- Ministral 3 8B 2512 (Mistral AI) is now listed at $0.15/M blended; no independent results yet. (details)
- Claude Opus 4.6 (no reasoning) (Anthropic) appeared on 6 benchmarks, entering the BenchLeader Index at #179. (details)
- Qwen3 6 (max) (Alibaba) appeared on 5 benchmarks, entering the BenchLeader Index at #60. (details)
- Gemini 3 Flash (thinking) (Google) appeared on 6 benchmarks, entering the BenchLeader Index at #66. (details)
- Grok 4.3 (no reasoning) (xAI) appeared on 6 benchmarks, entering the BenchLeader Index at #327. (details)
- Claude Opus 4.7 (thinking) (Anthropic) appeared on 1 benchmark. (details)
- KAT-Coder-Pro V2 (KwaiKAT) appeared on 5 benchmarks, entering the BenchLeader Index at #131. (details)
- Qwen3.5 Plus (thinking) (Alibaba) appeared on 6 benchmarks, entering the BenchLeader Index at #147. (details)
- Gemini 3 Pro (low) (Google) appeared on 5 benchmarks, entering the BenchLeader Index at #161. (details)
- Claude Sonnet 4.6 (low) (Anthropic) appeared on 6 benchmarks, entering the BenchLeader Index at #174. (details)
- Kimi K2.6 (no reasoning) (Moonshot AI) appeared on 5 benchmarks, entering the BenchLeader Index at #180. (details)
- GPT-5.1 (thinking) (OpenAI) appeared on 6 benchmarks, entering the BenchLeader Index at #181. (details)
- GLM-5 (no reasoning) (Zhipu AI) appeared on 5 benchmarks, entering the BenchLeader Index at #197. (details)
- Command A+ (Cohere) appeared on 6 benchmarks, entering the BenchLeader Index at #213. (details)
- Grok 3 Mini Fast (high) (xAI) appeared on 8 benchmarks, entering the BenchLeader Index at #222. (details)
- Qwen3.5 Omni Plus (Alibaba) appeared on 6 benchmarks, entering the BenchLeader Index at #229. (details)
- Kimi K2.5 (no reasoning) (Moonshot AI) appeared on 6 benchmarks, entering the BenchLeader Index at #241. (details)
- Gemini 2.5 Flash 09 (thinking) (Google) appeared on 14 benchmarks, entering the BenchLeader Index at #246. (details)
- Gemma 4 12B (no reasoning) (Google) appeared on 6 benchmarks, entering the BenchLeader Index at #390. (details)
- JT-35B-Flash (China Mobile) appeared on 5 benchmarks, entering the BenchLeader Index at #259. (details)
- Qwen3.5 27B (no reasoning) (Alibaba) appeared on 6 benchmarks, entering the BenchLeader Index at #262. (details)
- Grok 3 mini (thinking) (xAI) appeared on 5 benchmarks, entering the BenchLeader Index at #274. (details)
- Nemotron Cascade 2 30B A3B (NVIDIA) appeared on 5 benchmarks, entering the BenchLeader Index at #275. (details)
- K-EXAONE (no reasoning) (LG AI Research) appeared on 5 benchmarks, entering the BenchLeader Index at #401. (details)
- Nova 2.0 Lite (high) (Amazon) appeared on 6 benchmarks, entering the BenchLeader Index at #479. (details)
- MiMo-V2.5-Pro (no reasoning) (Xiaomi) appeared on 5 benchmarks, entering the BenchLeader Index at #284. (details)
- Gemma-4-31B-IT (no reasoning) (NVIDIA) appeared on 6 benchmarks, entering the BenchLeader Index at #285. (details)
- GLM-4.7 (no reasoning) (Zhipu AI) appeared on 5 benchmarks, entering the BenchLeader Index at #286. (details)
- Grok 3 Mini Fast (low) (xAI) appeared on 8 benchmarks, entering the BenchLeader Index at #288. (details)
- EXAONE 4.5 33B (no reasoning) (LG AI Research) appeared on 0 benchmarks. (details)
- North Mini Code (Cohere) appeared on 5 benchmarks, entering the BenchLeader Index at #301. (details)
- HyperNova 60B 2605 (high, based on gpt-oss-120b) (Multiverse Computing) appeared on 5 benchmarks, entering the BenchLeader Index at #308. (details)
- K2 Think V2 (low) (MBZUAI Institute of Foundation Models) appeared on 5 benchmarks, entering the BenchLeader Index at #441. (details)
- Gemini 2.5 Flash Lite 09 (Google) appeared on 14 benchmarks, entering the BenchLeader Index at #384. (details)
- nemotron-3-nano-30b-a3b (thinking) (NVIDIA) appeared on 5 benchmarks, entering the BenchLeader Index at #320. (details)
- Nova 2.0 Omni (medium) (Amazon) appeared on 6 benchmarks, entering the BenchLeader Index at #475. (details)
- Ernie 5.0 (thinking) (Baidu) appeared on 5 benchmarks, entering the BenchLeader Index at #328. (details)
- Ring 1t (Unknown) appeared on 6 benchmarks, entering the BenchLeader Index at #340. (details)
- Gemma 4 26B A4B (no reasoning) (Google) appeared on 6 benchmarks, entering the BenchLeader Index at #345. (details)
- Qwen3.5-4B (Alibaba) appeared on 8 benchmarks, entering the BenchLeader Index at #356. (details)
- Mistral Small 4 (no reasoning) (Mistral AI) appeared on 6 benchmarks, entering the BenchLeader Index at #466. (details)
- Gemini 2.5 Flash-Lite (thinking) (Google) appeared on 12 benchmarks, entering the BenchLeader Index at #363. (details)
- Nova 2.0 Pro Preview (Amazon) appeared on 5 benchmarks, entering the BenchLeader Index at #365. (details)
- DiffusionGemma 26B A4B (Google) appeared on 5 benchmarks, entering the BenchLeader Index at #369. (details)
- GPT-4.1 (high) (OpenAI) appeared on 8 benchmarks, entering the BenchLeader Index at #370. (details)
- Solar Open 100B (thinking) (Upstage) appeared on 5 benchmarks, entering the BenchLeader Index at #371. (details)
- Nemotron 3 Nano Omni 30B A3b (NVIDIA) appeared on 6 benchmarks, entering the BenchLeader Index at #377. (details)
- GPT-4.1 mini (high) (OpenAI) appeared on 8 benchmarks, entering the BenchLeader Index at #381. (details)
- Qwen3.5 Omni Flash (Alibaba) appeared on 6 benchmarks, entering the BenchLeader Index at #387. (details)
- Gemma 4 E4B (no reasoning) (Google) appeared on 6 benchmarks, entering the BenchLeader Index at #443. (details)
- GLM-4.6V (thinking) (Zhipu AI) appeared on 6 benchmarks, entering the BenchLeader Index at #395. (details)
- Ling 1t (Unknown) appeared on 6 benchmarks, entering the BenchLeader Index at #399. (details)
- MiniCPM5-1B (no reasoning) (OpenBMB) appeared on 5 benchmarks, entering the BenchLeader Index at #424. (details)
- Llama 3.3 Nemotron Super 49B v1.5 (thinking) (NVIDIA) appeared on 5 benchmarks, entering the BenchLeader Index at #476. (details)
- LFM2.5-8B-A1B (Liquid AI) appeared on 5 benchmarks, entering the BenchLeader Index at #419. (details)
- Qwen3 Vl 32B (Alibaba) appeared on 5 benchmarks, entering the BenchLeader Index at #420. (details)
- Hermes 4 - Llama-3.1 405B (thinking) (Nous Research) appeared on 5 benchmarks, entering the BenchLeader Index at #430. (details)
- Nemotron 3 Nano 4B (NVIDIA) appeared on 5 benchmarks, entering the BenchLeader Index at #425. (details)
- Tri-21B-Think (thinking) (Trillion Labs) appeared on 5 benchmarks, entering the BenchLeader Index at #427. (details)
- LongCat Flash Lite (LongCat) appeared on 5 benchmarks, entering the BenchLeader Index at #431. (details)
- Motif-2-12.7B (Motif Technologies) appeared on 5 benchmarks, entering the BenchLeader Index at #432. (details)
- HyperCLOVA X SEED Think (32B) (thinking) (Naver) appeared on 5 benchmarks, entering the BenchLeader Index at #434. (details)
- JT-MINI (China Mobile) appeared on 5 benchmarks, entering the BenchLeader Index at #435. (details)
- Falcon-H1R-7B (TII UAE) appeared on 5 benchmarks, entering the BenchLeader Index at #444. (details)
- Qwen3-VL 30B-A3B (Alibaba) appeared on 5 benchmarks, entering the BenchLeader Index at #445. (details)
- Step3 VL 10B (StepFun) appeared on 6 benchmarks, entering the BenchLeader Index at #447. (details)
- Mi:dm K 2.5 Pro (Korea Telecom) appeared on 5 benchmarks, entering the BenchLeader Index at #448. (details)
- Llama Nemotron Super 49B v1.5 (NVIDIA) appeared on 5 benchmarks, entering the BenchLeader Index at #458. (details)
- Jamba Reasoning 3B (thinking) (AI21 Labs) appeared on 5 benchmarks, entering the BenchLeader Index at #463. (details)
- Gemma 4 E2B (no reasoning) (Google) appeared on 6 benchmarks, entering the BenchLeader Index at #531. (details)
- Nemotron Nano 12B v2 VL (thinking) (NVIDIA) appeared on 6 benchmarks, entering the BenchLeader Index at #471. (details)
- nvidia-nemotron-nano-9b-v2 (thinking) (NVIDIA) appeared on 5 benchmarks, entering the BenchLeader Index at #472. (details)
- Devstral Small 2 (Mistral AI) appeared on 6 benchmarks, entering the BenchLeader Index at #474. (details)
- Qwen3 Omni 30B A3B Instruct (Alibaba) appeared on 6 benchmarks, entering the BenchLeader Index at #537. (details)
- Jamba 1.7 Large (AI21 Labs) appeared on 5 benchmarks, entering the BenchLeader Index at #483. (details)
- Qwen3 14B (thinking) (Alibaba) appeared on 5 benchmarks, entering the BenchLeader Index at #484. (details)
- Hermes 4 70B (Nous Research) appeared on 5 benchmarks, entering the BenchLeader Index at #514. (details)
- Nanbeige4.1-3B (Nanbeige) appeared on 5 benchmarks, entering the BenchLeader Index at #490. (details)
- GPT-4o (high) (OpenAI) appeared on 8 benchmarks, entering the BenchLeader Index at #493. (details)
- LFM2 24B A2B (Liquid AI) appeared on 5 benchmarks, entering the BenchLeader Index at #494. (details)
- Qwen3 32B (thinking) (Alibaba) appeared on 5 benchmarks, entering the BenchLeader Index at #497. (details)
- GLM-4.5V (thinking) (Zhipu AI) appeared on 6 benchmarks, entering the BenchLeader Index at #503. (details)
- Ministral 3 14B (Mistral AI) appeared on 6 benchmarks, entering the BenchLeader Index at #505. (details)
- Qwen3 Vl 8B (Alibaba) appeared on 6 benchmarks, entering the BenchLeader Index at #508. (details)
- Sarvam 105B (high) (Sarvam) appeared on 5 benchmarks, entering the BenchLeader Index at #509. (details)
- Granite 4.0 H Small (IBM) appeared on 5 benchmarks, entering the BenchLeader Index at #512. (details)
- Nova Micro (Amazon) appeared on 9 benchmarks, entering the BenchLeader Index at #513. (details)
- qwen3-4b-instruct-2507 (thinking) (Alibaba) appeared on 8 benchmarks, entering the BenchLeader Index at #515. (details)
- LFM2.5-1.2B-Instruct (thinking) (Liquid AI) appeared on 5 benchmarks, entering the BenchLeader Index at #534. (details)
- Olmo 3 7B Think (Allen Institute for AI) appeared on 5 benchmarks, entering the BenchLeader Index at #551. (details)
- Qwen3.5 2B (no reasoning) (Alibaba) appeared on 6 benchmarks, entering the BenchLeader Index at #558. (details)
- Ministral 3 8B (Mistral AI) appeared on 6 benchmarks, entering the BenchLeader Index at #526. (details)
- Jamba 1.7 Mini (AI21 Labs) appeared on 5 benchmarks, entering the BenchLeader Index at #527. (details)
- Qwen3 8B (thinking) (Alibaba) appeared on 5 benchmarks, entering the BenchLeader Index at #528. (details)
- Amazon Nova Lite (Amazon) appeared on 12 benchmarks, entering the BenchLeader Index at #529. (details)
- Granite 4.1 3B (IBM) appeared on 5 benchmarks, entering the BenchLeader Index at #532. (details)
- Gemma 3 270M (Google) appeared on 5 benchmarks, entering the BenchLeader Index at #538. (details)
- Reka Flash 3 (Reka AI) appeared on 5 benchmarks, entering the BenchLeader Index at #541. (details)
- Llama 4 Maverick Basic (Meta) appeared on 10 benchmarks, entering the BenchLeader Index at #542. (details)
- Llama Llama 3.3 70B Turbo (Meta) appeared on 6 benchmarks, entering the BenchLeader Index at #545. (details)
- Apertus 70B Instruct (Swiss AI Initiative) appeared on 5 benchmarks, entering the BenchLeader Index at #546. (details)
- Granite 4.0 H 1B (IBM) appeared on 5 benchmarks, entering the BenchLeader Index at #547. (details)
- Sarvam 30B (high) (Sarvam) appeared on 5 benchmarks, entering the BenchLeader Index at #549. (details)
- LFM2 2.6B (Liquid AI) appeared on 5 benchmarks, entering the BenchLeader Index at #550. (details)
- Qwen3-1.7B (thinking) (Alibaba) appeared on 5 benchmarks, entering the BenchLeader Index at #557. (details)
- Ministral 3 3B (Mistral AI) appeared on 6 benchmarks, entering the BenchLeader Index at #559. (details)
- Exaone 4.0 1.2B (thinking) (LG AI Research) appeared on 5 benchmarks, entering the BenchLeader Index at #566. (details)
- Apertus 8B Instruct (Swiss AI Initiative) appeared on 5 benchmarks, entering the BenchLeader Index at #565. (details)
- Granite 4.0 1B (IBM) appeared on 5 benchmarks, entering the BenchLeader Index at #567. (details)
- Qwen3 0 6B (thinking) (Qwen) appeared on 5 benchmarks, entering the BenchLeader Index at #568. (details)
- Tiny Aya Global (Cohere) appeared on 5 benchmarks, entering the BenchLeader Index at #576. (details)
- Granite 4.0 H 350M (IBM) appeared on 5 benchmarks, entering the BenchLeader Index at #577. (details)
- Llama 3.2 Instruct 11B (Vision) (Meta) appeared on 7 benchmarks, entering the BenchLeader Index at #579. (details)
- Molmo2-8B (Allen Institute for AI) appeared on 6 benchmarks, entering the BenchLeader Index at #583. (details)
- MiniCPM-V 4.6 1.3B (OpenBMB) appeared on 6 benchmarks, entering the BenchLeader Index at #584. (details)
- Qwen3.5 0.8B (no reasoning) (Alibaba) appeared on 6 benchmarks, entering the BenchLeader Index at #602. (details)
- Llama Llama 4 Scout 17B 16e (Meta) appeared on 8 benchmarks, entering the BenchLeader Index at #590. (details)
- LFM2.5-VL-1.6B (Liquid AI) appeared on 6 benchmarks, entering the BenchLeader Index at #599. (details)
- GPT-4o mini (high) (OpenAI) appeared on 7 benchmarks, entering the BenchLeader Index at #603. (details)
- GPT-4.1 nano (high) (OpenAI) appeared on 8 benchmarks, entering the BenchLeader Index at #607. (details)
- Jamba Large 1.6 (AI21 Labs) appeared on 8 benchmarks, entering the BenchLeader Index at #609. (details)
- Jamba Mini 1.6 (AI21 Labs) appeared on 8 benchmarks, entering the BenchLeader Index at #610. (details)
- Agnes 2.5 Pro Beta (Sapiens AI) appeared on 4 benchmarks. (details)
- Apodex 1.1 (Apodex) appeared on 4 benchmarks. (details)
- Apriel-v1.6-15B-Thinker (ServiceNow) appeared on 4 benchmarks. (details)
- Claude Sonnet 5 (medium) (Anthropic) appeared on 3 benchmarks. (details)
- Cogito v2.1 (thinking) (Deep Cogito) appeared on 4 benchmarks. (details)
- EXAONE 4.0 32B (LG AI Research) appeared on 4 benchmarks. (details)
- GPT 5.1 Instant (OpenAI) appeared on 4 benchmarks. (details)
- JT-4.1 Flash 236B A21B (China Mobile) appeared on 4 benchmarks. (details)
- Kimi Linear 48B A3B Instruct (Kimi) appeared on 4 benchmarks. (details)
- Langston Nim Nvidia Llama 3.3 Nemotron Super 49B v1 (Unknown) appeared on 4 benchmarks. (details)
- Langston Nim Nvidia Llama 3.3 Nemotron Super 49B v1 42e84561 (thinking) (Unknown) appeared on 4 benchmarks. (details)
- LFM2 8B A1B (Liquid AI) appeared on 4 benchmarks. (details)
- Llama Nemotron Super 49B v1.5 (thinking) (NVIDIA) appeared on 4 benchmarks. (details)
- Llama Nemotron Ultra 253B (thinking) (NVIDIA) appeared on 4 benchmarks. (details)
- Mistral Medium 3.5 (high) (Mistral AI) appeared on 4 benchmarks. (details)
- Nex-N2-Pro (Nex AGI) appeared on 4 benchmarks. (details)
- A.X-K2 (SK Telecom) appeared on 3 benchmarks. (details)
- Agnes 2.5 Pro Alpha (Sapiens AI) appeared on 3 benchmarks. (details)
- Celeris-1 (Celeris) appeared on 3 benchmarks. (details)
- G9v3-39A5B (AI9Stars) appeared on 3 benchmarks. (details)
- G9v3-3B (AI9Stars) appeared on 3 benchmarks. (details)
- K-EXAONE 2.0 (LG AI Research) appeared on 3 benchmarks. (details)
- K2 Horizon 375B A23B (MBZUAI Institute of Foundation Models) appeared on 3 benchmarks. (details)
- Kimi K3 (thinking) (Moonshot AI) appeared on 3 benchmarks. (details)
- LFM2.5-2.6B (Liquid AI) appeared on 3 benchmarks. (details)
- Ling 3.0 Tiny (InclusionAI) appeared on 3 benchmarks. (details)
- Llama 2 70B Steerlm (NVIDIA) appeared on 3 benchmarks. (details)
- Llama Meta Llama 3.1 70B Turbo (Meta) appeared on 3 benchmarks. (details)
- Llama Meta Llama 3.1 8B Turbo (Meta) appeared on 3 benchmarks. (details)
- LongCat 2.0 (LongCat) appeared on 3 benchmarks. (details)
- MiniCPM5-2B (OpenBMB) appeared on 3 benchmarks. (details)
- Motif 3 (Beta) (Motif Technologies) appeared on 3 benchmarks. (details)
- Muse Spark 1.3 (max) (Meta) appeared on 2 benchmarks. (details)
- Qed Nano (LM-Provers) appeared on 3 benchmarks. (details)
- Quasar 438B (max, based on GLM-5.2) (Multiverse Computing) appeared on 3 benchmarks. (details)
- Qwen-14B (Alibaba) appeared on 3 benchmarks. (details)
- Solar Open2 250B (Upstage) appeared on 3 benchmarks. (details)
- Zephyr Orpo 141B A35b (Hugging Face) appeared on 3 benchmarks. (details)
- GPT-3.5-turbo (high) (OpenAI) appeared on 2 benchmarks. (details)
- GPT-4 Turbo (high) (OpenAI) appeared on 2 benchmarks. (details)
- GPT-5.4 (thinking) (OpenAI) appeared on 2 benchmarks. (details)
- Llama Meta Llama 3.1 405B Turbo (Meta) appeared on 2 benchmarks. (details)
- Mistral Medium (latest) (Mistral AI) appeared on 2 benchmarks. (details)
- Nemotron Lightning 3.5 30B A3b (NVIDIA) appeared on 2 benchmarks. (details)
- Qwen Qwen2 5 72B Turbo (Alibaba) appeared on 2 benchmarks. (details)
- Qwen Qwen3 235B A22b (Alibaba) appeared on 2 benchmarks. (details)
- Afm 4 5B (Unknown) appeared on 1 benchmark. (details)
- Anubis 70B (Unknown) appeared on 1 benchmark. (details)
- Claude 37 Sonnet (thinking) (Anthropic) appeared on 1 benchmark. (details)
- Claude Fable 5 (thinking) (Anthropic) appeared on 1 benchmark. (details)
- Claude Haiku 3 (Anthropic) appeared on 1 benchmark. (details)
- Claude Haiku 4 (Anthropic) appeared on 1 benchmark. (details)
- Claude Opus 4.1 Anthropic (Anthropic) appeared on 1 benchmark. (details)
- Claude Opus 4.8 (thinking) (Anthropic) appeared on 1 benchmark. (details)
- Claude Opus 4.8 Claude Code (Anthropic) appeared on 1 benchmark. (details)
- Cogito v2 Deepseek 671B (Unknown) appeared on 1 benchmark. (details)
- Cogito v2 Llama 4 Scout (Unknown) appeared on 1 benchmark. (details)
- Cogito v2 Llama 405B (Unknown) appeared on 1 benchmark. (details)
- Cogito v2.1 (Deep Cogito) appeared on 1 benchmark. (details)
- DeepHermes 3 - Llama-3.1 8B (Nous Research) appeared on 1 benchmark. (details)
- DeepHermes 3 - Mistral 24B (Nous Research) appeared on 1 benchmark. (details)
- Gemini 1.0 Pro 002 (Google) appeared on 1 benchmark. (details)
- Gemini 2.5 Deep Think (thinking) (Google) appeared on 1 benchmark. (details)
- Gemini 3.1 Pro (thinking) (Google) appeared on 1 benchmark. (details)
- Gemini Flash 2.0 (Google) appeared on 1 benchmark. (details)
- Gemma 3 27B It Fast (Google) appeared on 1 benchmark. (details)
- Glm 4 32B (Zhipu AI) appeared on 1 benchmark. (details)
- GLM-5.2 (thinking) (Zhipu AI) appeared on 1 benchmark. (details)
- GPT 5 Agent (high) (OpenAI) appeared on 1 benchmark. (details)
- GPT 5.5 Codex (OpenAI) appeared on 1 benchmark. (details)
- GPT 5.5 Factory (OpenAI) appeared on 1 benchmark. (details)
- GPT-5 (thinking) (OpenAI) appeared on 1 benchmark. (details)
- GPT-5 mini (thinking) (OpenAI) appeared on 1 benchmark. (details)
- GPT-5.1-Codex (high) (OpenAI) appeared on 1 benchmark. (details)
- Grok 3 (thinking) (xAI) appeared on 1 benchmark. (details)
- Grok 4 Fast R (xAI) appeared on 1 benchmark. (details)
- Hermes 4 405B (thinking) (Nous Research) appeared on 1 benchmark. (details)
- Jamba Large 1.7 (AI21 Labs) appeared on 1 benchmark. (details)
- Jamba Mini 1.7 (AI21 Labs) appeared on 1 benchmark. (details)
- Kimi K2 (thinking) (Moonshot AI) appeared on 1 benchmark. (details)
- Llama 3 405B (Meta) appeared on 1 benchmark. (details)
- Llama 4 Maverick 17B (Meta) appeared on 1 benchmark. (details)
- Llama Llama 2 70B (Meta) appeared on 1 benchmark. (details)
- Mercury Coder (Inception) appeared on 1 benchmark. (details)
- Minimax 2.1 (MiniMax) appeared on 1 benchmark. (details)
- Openhands Lm 32B (All Hands AI) appeared on 1 benchmark. (details)
- Phi 4 Plus (thinking) (Microsoft) appeared on 1 benchmark. (details)
- Qwen Qwen2 5 7B Turbo (Alibaba) appeared on 1 benchmark. (details)
- Qwen3 235B A22B (no reasoning) (Alibaba) appeared on 1 benchmark. (details)
- Qwen3 32B Fast (thinking) (Alibaba) appeared on 1 benchmark. (details)
- Qwen3 5 397B (Alibaba) appeared on 1 benchmark. (details)
- Seed 1.6 (thinking) (ByteDance) appeared on 1 benchmark. (details)
- Sonar Pro (Perplexity) appeared on 1 benchmark. (details)
- Togethercomputer Llama 2 13B (Unknown) appeared on 1 benchmark. (details)
- Togethercomputer Llama 2 7B (Unknown) appeared on 1 benchmark. (details)
- Ui Tars 1.5 7B (ByteDance) appeared on 1 benchmark. (details)
- Valkyrie 49B v1 (Unknown) appeared on 1 benchmark. (details)
- Virtuoso Large (Unknown) appeared on 1 benchmark. (details)
- GPT-5.2 Codex (high) (OpenAI) appeared on 0 benchmarks. (details)
- Llama3-SWE-RL-70B (Meta) appeared on 0 benchmarks. (details)
- Muse Spark (thinking) (Meta) appeared on 0 benchmarks. (details)
- Qwen-1_8B (Alibaba) appeared on 0 benchmarks. (details)
- Qwen-7B (Alibaba) appeared on 0 benchmarks. (details)
- Qwen2.5-Max (Alibaba) appeared on 0 benchmarks. (details)
- riva-translate-4b-instruct-v1_1 (NVIDIA) appeared on 0 benchmarks. (details)
Price changes
- Llama 3-70B input price moved from $0.65/M to $0.51/M (22% down).
- Llama 3-70B output price moved from $2.75/M to $0.74/M (73% down).
- Llama 3.1 70B input price moved from $0.72/M to $0.56/M (22% down).
- Llama 3.1 70B output price moved from $0.72/M to $0.56/M (22% down).
- Nova Micro input price moved from $0.04/M to $0.04/M (12% down).
- Llama 3-8B input price moved from $0.05/M to $0.04/M (20% down).
- Llama 3-8B output price moved from $0.14/M to $0.04/M (71% down).
- Mixtral 8x7B input price moved from $0.45/M to $0.00/M (100% down).
- Mixtral 8x7B output price moved from $0.70/M to $0.00/M (100% down).
- Reka Flash 3 input price moved from $0.20/M to $0.10/M (50% down).
- Reka Flash 3 output price moved from $0.80/M to $0.20/M (75% down).
- Qwen3 235B A22B 2507 input price moved from $0.22/M to $0.20/M (9% down).
- Qwen3 235B A22B 2507 output price moved from $0.88/M to $0.60/M (32% down).
- Qwen3 235B A22B 2507 (thinking) input price moved from $0.22/M to $0.20/M (9% down).
- Qwen3 235B A22B 2507 (thinking) output price moved from $0.88/M to $0.60/M (32% down).
Top five as of 13 Sept 2026
- GPT-6 Astra — 70.5
- Claude Fable 5.1 — 70.1
- Claude Fable 5.1 — 69.1
- Claude Fable 5 — 68.8
- GPT-6 Astra — 68.7
Generated from the week's daily updates; every line links to the page where the numbers and their sources can be checked. Prices are list prices per million tokens, e.g. $0.510.