BenchLeader

MATH 500

Five hundred competition problems in algebra, probability and trigonometry. Run by Vals AI.

As of 19 Sept 2026, Gemini 3 Pro leads MATH 500 on BenchLeader with 96.4%, ahead of Grok 4 at 96.2%, across 57 model configurations with a published result.

Published by
Vals AI
Category
Maths
Index weight
Reference only
Models
57
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

MATH 500 is a 500-problem subset of the MATH competition dataset spanning algebra, counting, geometry and number theory, each with a short exact answer.

How it is scored

Accuracy, run by Vals AI at the stated reasoning effort.

What to keep in mind

Frontier models score above 95%, so it separates smaller models more than leaders, and the dataset is public.

57 of 57
#
1Gemini 3 ProhighGoogle96.4%61.12026-01-09
2Grok 4xAI96.2%58.52026-01-09
3GPT-5highOpenAI96.0%58.62026-01-09
4Claude Opus 4.1thinkingAnthropic95.4%57.22026-01-09
5Gemini 2.5 ProGoogle95.2%54.52026-01-09
6GPT-5 minihighOpenAI94.8%55.32026-01-09
7gpt-oss-120bOpenAIopen94.8%46.92026-01-09
8o3highOpenAI94.6%53.02026-01-09
9Qwen3 235B A22BAlibabaopen94.6%49.62026-01-09
10Grok 3 Mini FasthighxAI94.2%53.72026-01-09
11Kimi K2Moonshot AIopen94.2%51.02026-01-09
12o4-minihighOpenAI94.2%50.82026-01-09
13gpt-oss-20bOpenAIopen94.2%44.32026-01-09
14GLM-4.5Zhipu AIopen94.0%52.12026-01-09
15Claude Sonnet 4thinkingAnthropic93.8%54.02026-01-09
16GPT-5 nanohighOpenAI93.8%48.22026-01-09
17Claude Opus 4.1Anthropic93.0%51.92026-01-09
18DeepSeek R1DeepSeekopen92.2%48.72026-01-09
19Gemini 2.5 FlashthinkingGoogle91.8%48.42026-01-09
20o3-minihighOpenAI91.8%47.72026-01-09
21Gemini 2.5 FlashGoogle91.6%52.52026-01-09
22Claude 3.7 SonnetthinkingAnthropic91.6%51.72026-01-09
23Langston Nim Nvidia Llama 3.3 Nemotron Super 49B v1 42e84561thinkingNVIDIA91.4%2026-01-09
24Claude Opus 4Anthropic90.4%52.92026-01-09
25o1highOpenAI90.4%44.52026-01-09
26Claude Sonnet 4Anthropic90.3%49.72026-01-09
27Grok 3xAI89.8%51.02026-01-09
28MiniMax-M2MiniMaxopen89.0%55.12026-01-09
29Gemini 2.0 FlashGoogle89.0%43.72026-01-09
30GPT-4.1 minihighOpenAI88.0%46.32026-01-09
31GPT-4.1highOpenAI87.2%46.62026-01-09
32Mistral Medium 3Mistral AI87.0%45.52026-01-09
33Llama 4 Maverick BasicMeta85.2%39.42026-01-09
34Gemini 2.0 FlashthinkingGoogle84.6%45.22026-01-09
35Gemini 1.5 Pro 002Google82.8%44.92026-01-09
36DeepSeek V3DeepSeekopen80.4%44.82026-01-09
37GPT-4.1 nanohighOpenAI80.2%32.22026-01-09
38Llama Llama 4 Scout 17B 16eMeta79.2%36.42026-01-09
39Gemini 1.5 Flash 002Google78.8%40.82026-01-09
40Grok 2xAIopen78.4%42.22026-01-09
41Claude 3.7 SonnetAnthropic76.8%50.22026-01-09
42Command ACohereopen76.2%42.32026-01-09
43GPT-4ohighOpenAI75.2%41.62026-01-09
44Mistral Large 2Mistral AIopen74.4%41.02026-01-09
45Llama Llama 3.3 70B TurboMeta73.4%39.22026-01-09
46GPT-4o minihighOpenAI72.6%33.62026-01-09
47Claude 3.5 SonnetAnthropic72.4%46.62026-01-09
48Llama Meta Llama 3.1 405B TurboMeta71.4%2026-01-09
49Langston Nim Nvidia Llama 3.3 Nemotron Super 49B v1NVIDIA71.2%2026-01-09
50Mistral Small 2402Mistral AI70.6%32.02026-01-09
51Grok 3 Mini FastlowxAI70.2%50.62026-01-09
52Mistral Small 3.1Mistral AI68.4%34.62026-01-09
53Llama Meta Llama 3.1 70B TurboMeta65.0%2026-01-09
54Claude 3.5 HaikuAnthropic64.2%40.32026-01-09
55Jamba Large 1.6AI21 Labs54.8%26.02026-01-09
56Llama Meta Llama 3.1 8B TurboMeta44.4%2026-01-09
57Jamba Mini 1.6AI21 Labs25.4%21.82026-01-09

Cite as: BenchLeader, “MATH 500 leaderboard”, https://www.benchleader.com/benchmarks/vals_math500, data as of 19 Sept 2026.

MATH 500: questions

What does MATH 500 measure?
MATH 500 is a 500-problem subset of the MATH competition dataset spanning algebra, counting, geometry and number theory, each with a short exact answer. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads MATH 500?
Gemini 3 Pro leads MATH 500 with 96.4% as of 19 Sept 2026, ahead of Grok 4 at 96.2%.
How many models have MATH 500 results?
57 model configurations have a MATH 500 result on BenchLeader, all taken from Vals AI.
Who runs MATH 500 and how often is it updated?
MATH 500 is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does MATH 500 count toward the BenchLeader Index?
No. MATH 500 is shown for reference but left out of the composite index.