MMLU-Pro (HELM)
Stanford CRFM's own MMLU-Pro run. Shown for reference; Vals AI's run is in the index.
As of 12 Sept 2026, Gemini 3 Pro leads MMLU-Pro (HELM) on BenchLeader with 90.3%, ahead of Claude Opus 4 at 87.5%, across 68 model configurations with a published result.
- Published by
- HELM Capabilities
- Category
- Knowledge
- Index weight
- Reference only
- Models
- 68
- Data as of
- 12 Sept 2026
Stanford Center for Research on Foundation Models (crfm.stanford.edu), Apache-2.0 results.
What the test looks like
MMLU-Pro is a harder successor to MMLU: about 12,000 questions across 14 subjects with ten answer options each and more reasoning-heavy items, removing the trivial and noisy questions of the original.
How it is scored
Accuracy, run by Vals AI.
What to keep in mind
Stanford CRFM's own run of MMLU-Pro in HELM's fixed harness. Shown for reference; Vals AI's run is counted in the index. See also the MMLU-Pro board.
- 1Gemini 3 Pro90.3%
- 2Claude Opus 487.5%
- 3Claude Sonnet 4.586.9%
- 4GPT-586.3%
- 5Gemini 2.5 Pro86.3%
- 6o385.9%
- 7Claude Opus 485.9%
- 8Grok 485.1%
- 9Qwen3 235B A22B 250784.4%
- 10Claude Sonnet 484.3%
- 11Claude Sonnet 484.3%
- 12GPT-5 mini83.5%
- 13o4-mini82.0%
- 14Kimi K281.9%
- 15Qwen3 235B A22B81.7%
| # | |||||
|---|---|---|---|---|---|
| 1 | 90.3% | – | 62.0 | – | |
| 2 | 87.5% | – | 56.1 | – | |
| 3 | 86.9% | 20250929 | 53.7 | – | |
| 4 | 86.3% | – | 60.8 | – | |
| 5 | 86.3% | – | 54.5 | – | |
| 6 | 85.9% | – | 61.6 | – | |
| 7 | 85.9% | 20250514 | 52.2 | – | |
| 8 | 85.1% | 0709 | 58.6 | – | |
| 9 | 84.4% | Qwen3 235B A22B Instruct 2507 FP8 | 54.1 | – | |
| 10 | 84.3% | – | 55.1 | – | |
| 11 | 84.3% | 20250514 | 50.5 | – | |
| 12 | 83.5% | – | 57.9 | – | |
| 13 | 82.0% | – | 57.2 | – | |
| 14 | 81.9% | Kimi K2 Instruct | 51.2 | – | |
| 15 | 81.7% | Qwen3 235B A22B FP8 Throughput | 44.8 | – | |
| 16 | 81.1% | – | 48.8 | – | |
| 17 | 81.0% | Llama 4 Maverick (17Bx128E) Instruct FP8 | 44.2 | – | |
| 18 | 80.4% | Palmyra X5 | – | – | |
| 19 | 79.9% | – | 51.8 | – | |
| 20 | 79.5% | – | 48.8 | – | |
| 21 | 79.3% | – | 49.9 | – | |
| 22 | 78.8% | – | 51.0 | – | |
| 23 | 78.6% | – | 51.2 | – | |
| 24 | 78.4% | 20250219 | 50.2 | – | |
| 25 | 78.3% | – | 46.3 | – | |
| 26 | 77.8% | – | 52.5 | – | |
| 27 | 77.7% | 20251001 | 48.6 | – | |
| 28 | 77.7% | 20241022 | 46.3 | – | |
| 29 | 76.2% | – | 47.0 | – | |
| 30 | 74.2% | Llama 4 Scout (17Bx16E) Instruct | 40.4 | – | |
| 31 | 74.0% | – | 46.0 | – | |
| 32 | 73.7% | Gemini 1.5 Pro (002) | 44.8 | – | |
| 33 | 73.7% | Gemini 2.0 Flash | 43.5 | – | |
| 34 | 72.6% | – | 45.8 | – | |
| 35 | 72.3% | DeepSeek v3 | 45.2 | – | |
| 36 | 72.3% | Llama 3.1 Instruct Turbo (405B) | 43.1 | – | |
| 37 | 72.0% | – | 47.7 | – | |
| 38 | 71.3% | – | 43.8 | – | |
| 39 | 67.8% | Gemini 1.5 Flash (002) | 40.8 | – | |
| 40 | 67.3% | – | 41.1 | – | |
| 41 | 65.7% | – | – | – | |
| 42 | 65.3% | Llama 3.1 Instruct Turbo (70B) | 41.1 | – | |
| 43 | 63.9% | – | 50.0 | – | |
| 44 | 63.1% | Qwen2.5 Instruct Turbo (72B) | 42.7 | – | |
| 45 | 61.0% | 2503 | 40.3 | – | |
| 46 | 60.5% | 20241022 | 40.0 | – | |
| 47 | 60.3% | – | 36.6 | – | |
| 48 | 60.0% | – | 39.4 | – | |
| 49 | 59.9% | 2411 | 40.8 | – | |
| 50 | 59.1% | – | – | – | |
| 51 | 57.9% | – | 59.4 | – | |
| 52 | 56.9% | IBM Granite 4.0 Small | – | – | |
| 53 | 55.0% | – | 39.1 | – | |
| 54 | 53.9% | Qwen2.5 Instruct Turbo (7B) | 40.1 | – | |
| 55 | 53.7% | – | 43.8 | – | |
| 56 | 51.1% | – | 39.8 | – | |
| 57 | 46.0% | Mixtral Instruct (8x22B) | 37.6 | – | |
| 58 | 41.4% | OLMo 2 32B Instruct March 2025 | – | – | |
| 59 | 41.1% | – | – | – | |
| 60 | 40.6% | Llama 3.1 Instruct Turbo (8B) | 36.3 | – |
Cite as: BenchLeader, “MMLU-Pro (HELM) leaderboard”, https://www.benchleader.com/benchmarks/helm_mmlu_pro, data as of 12 Sept 2026.
MMLU-Pro (HELM): questions
- What does MMLU-Pro (HELM) measure?
- MMLU-Pro is a harder successor to MMLU: about 12,000 questions across 14 subjects with ten answer options each and more reasoning-heavy items, removing the trivial and noisy questions of the original. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads MMLU-Pro (HELM)?
- Gemini 3 Pro leads MMLU-Pro (HELM) with 90.3% as of 12 Sept 2026, ahead of Claude Opus 4 at 87.5%.
- How many models have MMLU-Pro (HELM) results?
- 68 model configurations have a MMLU-Pro (HELM) result on BenchLeader, all taken from HELM Capabilities.
- Who runs MMLU-Pro (HELM) and how often is it updated?
- MMLU-Pro (HELM) is published by HELM Capabilities. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does MMLU-Pro (HELM) count toward the BenchLeader Index?
- No. MMLU-Pro (HELM) is shown for reference but left out of the composite index.