BenchLeader

MMMU-Pro (Vals)

Vals AI's own run of MMMU-Pro. Run by Vals AI.

As of 19 Sept 2026, Claude Fable 5.1 leads MMMU-Pro (Vals) on BenchLeader with 90.6%, ahead of Claude Opus 5 at 89.9%, across 90 model configurations with a published result.

Published by
Vals AI
Category
Multimodal
Index weight
Reference only
Models
90
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

Vals AI's run of the harder ten-option MMMU-Pro multimodal exam.

How it is scored

Accuracy, run by Vals AI.

What to keep in mind

Shown next to the Artificial Analysis run and the official board.

90 of 90
#
1Claude Fable 5.1Anthropic90.6%64.52026-09-01
2Claude Opus 5Anthropic89.9%66.72026-09-01
3Claude Fable 5Anthropic89.3%68.32026-09-01
4Gemini 3.8 FlashhighGoogle89.1%64.52026-09-01
5Gemini 3.7 FlashhighGoogle89.0%64.62026-09-01
6GPT-5.6 SolmaxOpenAI88.8%68.82026-09-01
7Gemini 3.6 FlashhighGoogle88.4%61.72026-09-01
8GPT-5.5xhighOpenAI88.3%67.62026-09-01
9Gemini 3.5 FlashhighGoogle88.3%63.62026-09-01
10Gemini 3.1 ProhighGoogle88.2%60.72026-09-01
11Kimi K3Moonshot AIopen ↗88.2%64.22026-09-01
12Qwen3 8maxAlibaba88.0%66.52026-09-01
13Gemini 3 FlashhighGoogle87.6%58.72026-09-01
14GPT-5.4xhighOpenAI87.5%65.42026-09-01
15Gemini 3 ProhighGoogle87.5%61.12026-09-01
16Muse SparkMeta87.4%65.72026-09-01
17GPT-5.2xhighOpenAI86.7%62.02026-09-01
18Claude Opus 4.8Anthropic86.6%61.92026-09-01
19Muse Spark 1.1xhighMeta86.6%61.82026-09-01
20GPT-5.6 TerramaxOpenAI86.5%65.02026-09-01
21Kimi K2.6Moonshot AIopen ↗86.3%60.52026-09-01
22Muse Spark 1.2xhighMeta86.1%64.12026-09-01
23GLM-5.3-FlashmaxZhipu AIopen ↗86.0%55.32026-09-01
24Claude Opus 4.7Anthropic85.5%64.52026-09-01
25GPT-5.6 LunamaxOpenAI85.0%60.22026-09-01
26Kimi K2.5thinkingMoonshot AIopen84.3%58.52026-09-01
27Qwen3.6 PlusAlibabaopen84.2%58.52026-09-01
28Qwen3.8 27BxhighAlibabaopen ↗83.9%59.22026-09-01
29Claude Opus 4.6thinkingAnthropic83.9%62.52026-09-01
30Claude Sonnet 4.6Anthropic83.6%59.02026-09-01
31Gemini 3.5 Flash LitehighGoogle83.6%44.32026-09-01
32Grok 4.20thinkingxAI83.5%59.92026-09-01
33GPT-5.1highOpenAI83.2%58.72026-09-01
34Grok 4.3xAI83.1%50.52026-09-01
35Claude Sonnet 5Anthropic83.0%57.22026-09-01
36Claude Opus 4.5thinkingAnthropic83.0%60.82026-09-01
37Gemini 3.1 Flash LitehighGoogle82.5%49.12026-09-01
38Qwen3.5 FlashAlibaba81.9%52.52026-09-01
39GPT-5highOpenAI81.5%58.62026-09-01
40Gemini 2.5 ProGoogle81.3%54.52026-09-01
41MiniMax-M3MiniMaxopen ↗81.2%57.52026-09-01
42Claude Opus 4.5Anthropic81.1%58.22026-09-01
43Gemini 2.5 Flash 09 2025thinkingGoogle80.8%51.62026-09-01
44o3highOpenAI80.4%53.02026-09-01
45o4-minihighOpenAI79.7%50.82026-09-01
46Gemini 2.5 Flash 09 2025Google79.5%52.12026-09-01
47Claude Sonnet 4.5thinkingAnthropic79.3%53.42026-09-01
48GPT-5.4 minixhighOpenAI79.3%55.72026-09-01
49GPT-5 minihighOpenAI78.9%55.32026-09-01
50Claude Opus 4.1thinkingAnthropic77.5%57.22026-09-01
51o1highOpenAI77.4%44.52026-09-01
52Grok 4xAI76.3%58.52026-09-01
53Gemini 2.5 Flash Lite 09thinkingGoogle75.4%48.32026-09-01
54Claude 3.7 SonnetthinkingAnthropic75.1%51.72026-09-01
55Claude Sonnet 4thinkingAnthropic74.9%54.02026-09-01
56Claude Opus 4.1Anthropic73.7%51.92026-09-01
57GPT-5.4 nanohighOpenAI73.6%49.12026-09-01
58Claude Opus 4Anthropic73.3%52.92026-09-01
59Grok 4 FastthinkingxAI72.8%52.62026-09-01
60Grok 4.1thinkingxAI72.7%56.02026-09-01

Cite as: BenchLeader, “MMMU-Pro (Vals) leaderboard”, https://www.benchleader.com/benchmarks/vals_mmmu_pro, data as of 19 Sept 2026.

MMMU-Pro (Vals): questions

What does MMMU-Pro (Vals) measure?
Vals AI's run of the harder ten-option MMMU-Pro multimodal exam. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads MMMU-Pro (Vals)?
Claude Fable 5.1 leads MMMU-Pro (Vals) with 90.6% as of 19 Sept 2026, ahead of Claude Opus 5 at 89.9%.
How many models have MMMU-Pro (Vals) results?
90 model configurations have a MMMU-Pro (Vals) result on BenchLeader, all taken from Vals AI.
Who runs MMMU-Pro (Vals) and how often is it updated?
MMMU-Pro (Vals) is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does MMMU-Pro (Vals) count toward the BenchLeader Index?
No. MMMU-Pro (Vals) is shown for reference but left out of the composite index.