BenchLeader

MMLU

Fifty-seven-subject multiple-choice exam of world knowledge. A 2020 benchmark, saturated for current models; kept for history.

As of 19 Sept 2026, GPT-4o leads MMLU on BenchLeader with 88.1%, ahead of Claude 3.5 Sonnet at 87.3%, across 121 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Knowledge
Index weight
Reference only
Models
121
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Fifty-seven-subject multiple-choice exam of world knowledge.

How it is scored

Accuracy, as compiled by Epoch AI from published results.

What to keep in mind

Saturated: frontier models score near the ceiling, so it separates only older or smaller models, and results are largely self-reported by developers.

121 of 121
#
1GPT-4oOpenAI88.1%44.32024-11-20
2Claude 3.5 SonnetAnthropic87.3%46.62024-10-22
3DeepSeek V3DeepSeekopen87.2%44.82024-12-26
4Gemini 1.5 Pro 002Google86.9%44.92024-09-24
5GPT 4 0314OpenAI86.4%43.92023-03-14
6Llama 3.3 70BMetaopen86.3%41.42024-12-06
7Gemini 1.5 Pro 001Google85.9%45.12024-05-14
8Qwen2.5 72BAlibabaopen85.3%42.72024-09-19
9Phi-4Microsoftopen84.8%39.62024-12-12
10Claude 3 OpusAnthropic84.6%39.62024-02-29
11Llama 3.1 405BMetaopen84.5%43.52024-07-23
12Gemini 1.5 Pro 001 Feb24Google82.7%2024-02-15
13Qwen2-72BAlibabaopen ↗82.4%41.32024-06-07
14GPT 4 0613OpenAI82.4%40.22023-06-13
15Nova ProAmazon82.0%41.42024-12-03
16GPT-4o miniOpenAI81.8%36.22024-07-18
17GPT-4 TurboOpenAI81.3%41.52024-04-09
18Llama 3.2 90BMetaopen ↗80.3%37.62024-09-24
19Llama 3.1 70BMetaopen80.1%41.32024-07-23
20Mistral Large 2Mistral AIopen80.0%41.02024-07-24
21Qwen2.5 14B InstructAlibabaopen ↗79.9%2024-09-19
22Gemini 2.0 FlashGoogle79.7%43.72024-12-11
23Llama 3-70BMetaopen79.3%39.22024-04-18
24Yi-Large01.AI79.3%2024-05-13
25Qwen2.5-Coder 32B InstructAlibabaopen ↗79.1%42.62024-09-18
26Claude 2Anthropic78.5%2023-07-11
27Deepseek v2DeepSeekopen ↗78.4%2024-05-07
28phi-3-medium 14BmediumMicrosoft78.0%2024-04-23
29Gemini 1.5 Flash 001Google77.9%36.72024-05-23
30Mixtral 8x22B InstructMistral AIopen ↗77.8%37.72024-04-17
31Gemini 1.5 Flash 0514Google77.8%2024-05-14
32Nova LiteAmazon77.0%39.52024-12-03
33Claude 1.3Anthropic77.0%2023-04-18
34Yi-34B01.AI76.3%35.52023-11-02
35Claude 3 SonnetAnthropic75.9%38.62024-02-29
36Gemma 2 27BGoogleopen ↗75.7%41.42024-06-24
37Phi 3 SmallMicrosoft75.7%2024-04-23
38Qwen2.5-Coder-14BAlibaba75.2%2024-09-18
39Qwen1.5-32BAlibaba74.4%40.32024-02-04
40Claude 3.5 HaikuAnthropic74.3%40.32024-10-22
41Gemini 1.5 Flash 002Google73.9%40.82024-09-24
42Claude 3 HaikuAnthropic73.8%37.62024-03-07
43Claude 2.1Anthropic73.5%2023-11-21
44Claude InstantAnthropic73.4%
45Claude InstantAnthropic73.2%2023-08-09
46Qwen2.5 7B InstructAlibaba72.9%40.12024-09-19
47Inflection-1Inflection AI72.7%2023-06-22
48Gemma 2 9BGoogle72.1%39.52024-06-24
49GPT 3.5 Turbo 1106OpenAI71.4%39.22023-11-06
50Nova MicroAmazon70.8%39.92024-12-03
51Mixtral 8x7BMistral AIopen70.6%32.12023-12-11
52Falcon-180BTII70.6%2023-09-06
53Gemini 1.0 ProGoogle70.0%2023-12-13
54davinci-002OpenAI70.0%2022-03-15
55Llama 2-70BMetaopen ↗69.9%34.92023-07-18
56Command R+Cohere69.4%2024-08-30
57Palm 540BGoogle69.3%2022-04-04
58GPT-3.5 Turbo (older v0613)OpenAI68.9%2023-06-13
59Mistral LargeMistral AI68.8%41.32024-02-26
60Phi 3 MiniMicrosoftopen68.8%36.22024-04-23

Cite as: BenchLeader, “MMLU leaderboard”, https://www.benchleader.com/benchmarks/mmlu, data as of 19 Sept 2026.

MMLU: questions

What does MMLU measure?
Fifty-seven-subject multiple-choice exam of world knowledge. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads MMLU?
GPT-4o leads MMLU with 88.1% as of 19 Sept 2026, ahead of Claude 3.5 Sonnet at 87.3%.
How many models have MMLU results?
121 model configurations have a MMLU result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs MMLU and how often is it updated?
MMLU is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does MMLU count toward the BenchLeader Index?
No. MMLU is shown for reference but left out of the composite index.