BenchLeader

MMLU-Pro (HELM)

Stanford CRFM's own MMLU-Pro run. Shown for reference; Vals AI's run is in the index.

As of 12 Sept 2026, Gemini 3 Pro leads MMLU-Pro (HELM) on BenchLeader with 90.3%, ahead of Claude Opus 4 at 87.5%, across 68 model configurations with a published result.

Published by
HELM Capabilities
Category
Knowledge
Index weight
Reference only
Models
68
Data as of
12 Sept 2026

Stanford Center for Research on Foundation Models (crfm.stanford.edu), Apache-2.0 results.

What the test looks like

MMLU-Pro is a harder successor to MMLU: about 12,000 questions across 14 subjects with ten answer options each and more reasoning-heavy items, removing the trivial and noisy questions of the original.

How it is scored

Accuracy, run by Vals AI.

What to keep in mind

Stanford CRFM's own run of MMLU-Pro in HELM's fixed harness. Shown for reference; Vals AI's run is counted in the index. See also the MMLU-Pro board.

68 of 68
#
1Gemini 3 ProGoogle90.3%62.0
2Claude Opus 4Anthropic87.5%56.1
3Claude Sonnet 4.5Anthropic86.9%2025092953.7
4GPT-5OpenAI86.3%60.8
5Gemini 2.5 ProGoogle86.3%54.5
6o3OpenAI85.9%61.6
7Claude Opus 4Anthropic85.9%2025051452.2
8Grok 4xAI85.1%070958.6
9Qwen3 235B A22B 2507Alibabaopen84.4%Qwen3 235B A22B Instruct 2507 FP854.1
10Claude Sonnet 4Anthropic84.3%55.1
11Claude Sonnet 4Anthropic84.3%2025051450.5
12GPT-5 miniOpenAI83.5%57.9
13o4-miniOpenAI82.0%57.2
14Kimi K2Moonshot AIopen81.9%Kimi K2 Instruct51.2
15Qwen3 235B A22BAlibabaopen81.7%Qwen3 235B A22B FP8 Throughput44.8
16GPT-4.1OpenAI81.1%48.8
17Llama 4 MaverickMetaopen81.0%Llama 4 Maverick (17Bx128E) Instruct FP844.2
18Palmyra X5Writer80.4%Palmyra X5
19Grok 3 minixAI79.9%51.8
20gpt-oss-120bOpenAIopen79.5%48.8
21DeepSeek R1 0528DeepSeekopen79.3%49.9
22Grok 3xAI78.8%51.0
23Qwen3-Next 80B A3BAlibabaopen78.6%51.2
24Claude 3.7 SonnetAnthropic78.4%2025021950.2
25GPT-4.1 miniOpenAI78.3%46.3
26GPT-5 nanoOpenAI77.8%52.5
27Claude Haiku 4.5Anthropic77.7%2025100148.6
28Claude 3.5 SonnetAnthropic77.7%2024102246.3
29GLM-4.5-AirZhipu AIopen76.2%47.0
30Llama 4 ScoutMetaopen74.2%Llama 4 Scout (17Bx16E) Instruct40.4
31gpt-oss-20bOpenAIopen74.0%46.0
32Gemini 1.5 Pro 002Google73.7%Gemini 1.5 Pro (002)44.8
33Gemini 2.0 FlashGoogle73.7%Gemini 2.0 Flash43.5
34Nova PremierAmazon72.6%45.8
35DeepSeek V3DeepSeekopen72.3%DeepSeek v345.2
36Llama 3.1 405BMetaopen72.3%Llama 3.1 Instruct Turbo (405B)43.1
37Gemini 2.0 Flash-LiteGoogle72.0%47.7
38GPT-4oOpenAI71.3%43.8
39Gemini 1.5 Flash 002Google67.8%Gemini 1.5 Flash (002)40.8
40Nova ProAmazon67.3%41.1
41Palmyra X 004Writer65.7%
42Llama 3.1 70BMetaopen65.3%Llama 3.1 Instruct Turbo (70B)41.1
43Gemini 2.5 FlashGoogle63.9%50.0
44Qwen2.5 72BAlibabaopen63.1%Qwen2.5 Instruct Turbo (72B)42.7
45Mistral Small 3.1Mistral AIopen61.0%250340.3
46Claude 3.5 HaikuAnthropic60.5%2024102240.0
47GPT-4o miniOpenAI60.3%36.6
48Amazon Nova LiteAmazon60.0%39.4
49Mistral Large 2Mistral AIopen59.9%241140.8
50Palmyra FinWriter59.1%
51GPT-5.1OpenAI57.9%59.4
52Granite 4.0 SmallIBM56.9%IBM Granite 4.0 Small
53GPT-4.1 nanoOpenAI55.0%39.1
54Qwen2.5 7B InstructAlibabaopen53.9%Qwen2.5 Instruct Turbo (7B)40.1
55Gemini 2.5 Flash-LiteGoogle53.7%43.8
56Nova MicroAmazon51.1%39.8
57Mixtral 8x22B InstructMistral AIopen46.0%Mixtral Instruct (8x22B)37.6
58Olmo 2 32B March 2025Ai241.4%OLMo 2 32B Instruct March 2025
59Palmyra MedWriter41.1%
60Llama 3.1 8BMetaopen40.6%Llama 3.1 Instruct Turbo (8B)36.3

Cite as: BenchLeader, “MMLU-Pro (HELM) leaderboard”, https://www.benchleader.com/benchmarks/helm_mmlu_pro, data as of 12 Sept 2026.

MMLU-Pro (HELM): questions

What does MMLU-Pro (HELM) measure?
MMLU-Pro is a harder successor to MMLU: about 12,000 questions across 14 subjects with ten answer options each and more reasoning-heavy items, removing the trivial and noisy questions of the original. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads MMLU-Pro (HELM)?
Gemini 3 Pro leads MMLU-Pro (HELM) with 90.3% as of 12 Sept 2026, ahead of Claude Opus 4 at 87.5%.
How many models have MMLU-Pro (HELM) results?
68 model configurations have a MMLU-Pro (HELM) result on BenchLeader, all taken from HELM Capabilities.
Who runs MMLU-Pro (HELM) and how often is it updated?
MMLU-Pro (HELM) is published by HELM Capabilities. BenchLeader re-reads the published results every morning and records the date each result was published.
Does MMLU-Pro (HELM) count toward the BenchLeader Index?
No. MMLU-Pro (HELM) is shown for reference but left out of the composite index.