MMLU
Fifty-seven-subject multiple-choice exam of world knowledge. A 2020 benchmark, saturated for current models; kept for history.
As of 19 Sept 2026, GPT-4o leads MMLU on BenchLeader with 88.1%, ahead of Claude 3.5 Sonnet at 87.3%, across 121 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Knowledge
- Index weight
- Reference only
- Models
- 121
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Fifty-seven-subject multiple-choice exam of world knowledge.
How it is scored
Accuracy, as compiled by Epoch AI from published results.
What to keep in mind
Saturated: frontier models score near the ceiling, so it separates only older or smaller models, and results are largely self-reported by developers.
- 1GPT-4o88.1%
- 2Claude 3.5 Sonnet87.3%
- 3DeepSeek V387.2%
- 4Gemini 1.5 Pro 00286.9%
- 5GPT 4 031486.4%
- 6Llama 3.3 70B86.3%
- 7Gemini 1.5 Pro 00185.9%
- 8Qwen2.5 72B85.3%
- 9Phi-484.8%
- 10Claude 3 Opus84.6%
- 11Llama 3.1 405B84.5%
- 12Gemini 1.5 Pro 001 Feb2482.7%
- 13Qwen2-72B82.4%
- 14GPT 4 061382.4%
- 15Nova Pro82.0%
| # | ||||
|---|---|---|---|---|
| 1 | 88.1% | 44.3 | 2024-11-20 | |
| 2 | 87.3% | 46.6 | 2024-10-22 | |
| 3 | 87.2% | 44.8 | 2024-12-26 | |
| 4 | 86.9% | 44.9 | 2024-09-24 | |
| 5 | 86.4% | 43.9 | 2023-03-14 | |
| 6 | 86.3% | 41.4 | 2024-12-06 | |
| 7 | 85.9% | 45.1 | 2024-05-14 | |
| 8 | 85.3% | 42.7 | 2024-09-19 | |
| 9 | 84.8% | 39.6 | 2024-12-12 | |
| 10 | 84.6% | 39.6 | 2024-02-29 | |
| 11 | 84.5% | 43.5 | 2024-07-23 | |
| 12 | 82.7% | – | 2024-02-15 | |
| 13 | 82.4% | 41.3 | 2024-06-07 | |
| 14 | 82.4% | 40.2 | 2023-06-13 | |
| 15 | 82.0% | 41.4 | 2024-12-03 | |
| 16 | 81.8% | 36.2 | 2024-07-18 | |
| 17 | 81.3% | 41.5 | 2024-04-09 | |
| 18 | 80.3% | 37.6 | 2024-09-24 | |
| 19 | 80.1% | 41.3 | 2024-07-23 | |
| 20 | 80.0% | 41.0 | 2024-07-24 | |
| 21 | 79.9% | – | 2024-09-19 | |
| 22 | 79.7% | 43.7 | 2024-12-11 | |
| 23 | 79.3% | 39.2 | 2024-04-18 | |
| 24 | 79.3% | – | 2024-05-13 | |
| 25 | 79.1% | 42.6 | 2024-09-18 | |
| 26 | 78.5% | – | 2023-07-11 | |
| 27 | 78.4% | – | 2024-05-07 | |
| 28 | 78.0% | – | 2024-04-23 | |
| 29 | 77.9% | 36.7 | 2024-05-23 | |
| 30 | 77.8% | 37.7 | 2024-04-17 | |
| 31 | 77.8% | – | 2024-05-14 | |
| 32 | 77.0% | 39.5 | 2024-12-03 | |
| 33 | 77.0% | – | 2023-04-18 | |
| 34 | 76.3% | 35.5 | 2023-11-02 | |
| 35 | 75.9% | 38.6 | 2024-02-29 | |
| 36 | 75.7% | 41.4 | 2024-06-24 | |
| 37 | 75.7% | – | 2024-04-23 | |
| 38 | 75.2% | – | 2024-09-18 | |
| 39 | 74.4% | 40.3 | 2024-02-04 | |
| 40 | 74.3% | 40.3 | 2024-10-22 | |
| 41 | 73.9% | 40.8 | 2024-09-24 | |
| 42 | 73.8% | 37.6 | 2024-03-07 | |
| 43 | 73.5% | – | 2023-11-21 | |
| 44 | 73.4% | – | – | |
| 45 | 73.2% | – | 2023-08-09 | |
| 46 | 72.9% | 40.1 | 2024-09-19 | |
| 47 | Inflection-1Inflection AI | 72.7% | – | 2023-06-22 |
| 48 | 72.1% | 39.5 | 2024-06-24 | |
| 49 | 71.4% | 39.2 | 2023-11-06 | |
| 50 | 70.8% | 39.9 | 2024-12-03 | |
| 51 | 70.6% | 32.1 | 2023-12-11 | |
| 52 | 70.6% | – | 2023-09-06 | |
| 53 | 70.0% | – | 2023-12-13 | |
| 54 | 70.0% | – | 2022-03-15 | |
| 55 | 69.9% | 34.9 | 2023-07-18 | |
| 56 | 69.4% | – | 2024-08-30 | |
| 57 | 69.3% | – | 2022-04-04 | |
| 58 | 68.9% | – | 2023-06-13 | |
| 59 | 68.8% | 41.3 | 2024-02-26 | |
| 60 | 68.8% | 36.2 | 2024-04-23 |
Cite as: BenchLeader, “MMLU leaderboard”, https://www.benchleader.com/benchmarks/mmlu, data as of 19 Sept 2026.
MMLU: questions
- What does MMLU measure?
- Fifty-seven-subject multiple-choice exam of world knowledge. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads MMLU?
- GPT-4o leads MMLU with 88.1% as of 19 Sept 2026, ahead of Claude 3.5 Sonnet at 87.3%.
- How many models have MMLU results?
- 121 model configurations have a MMLU result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs MMLU and how often is it updated?
- MMLU is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does MMLU count toward the BenchLeader Index?
- No. MMLU is shown for reference but left out of the composite index.