MATH 500
Five hundred competition problems in algebra, probability and trigonometry. Run by Vals AI.
As of 19 Sept 2026, Gemini 3 Pro leads MATH 500 on BenchLeader with 96.4%, ahead of Grok 4 at 96.2%, across 57 model configurations with a published result.
- Published by
- Vals AI
- Category
- Maths
- Index weight
- Reference only
- Models
- 57
- Data as of
- 19 Sept 2026
Vals AI (vals.ai).
What the test looks like
MATH 500 is a 500-problem subset of the MATH competition dataset spanning algebra, counting, geometry and number theory, each with a short exact answer.
How it is scored
Accuracy, run by Vals AI at the stated reasoning effort.
What to keep in mind
Frontier models score above 95%, so it separates smaller models more than leaders, and the dataset is public.
- 1Gemini 3 Pro (high)96.4%
- 2Grok 496.2%
- 3GPT-5 (high)96.0%
- 4Claude Opus 4.1 (thinking)95.4%
- 5Gemini 2.5 Pro95.2%
- 6GPT-5 mini (high)94.8%
- 7gpt-oss-120b94.8%
- 8o3 (high)94.6%
- 9Qwen3 235B A22B94.6%
- 10Grok 3 Mini Fast (high)94.2%
- 11Kimi K294.2%
- 12o4-mini (high)94.2%
- 13gpt-oss-20b94.2%
- 14GLM-4.594.0%
- 15Claude Sonnet 4 (thinking)93.8%
57 of 57
| # | ||||
|---|---|---|---|---|
| 1 | 96.4% | 61.1 | 2026-01-09 | |
| 2 | 96.2% | 58.5 | 2026-01-09 | |
| 3 | 96.0% | 58.6 | 2026-01-09 | |
| 4 | 95.4% | 57.2 | 2026-01-09 | |
| 5 | 95.2% | 54.5 | 2026-01-09 | |
| 6 | 94.8% | 55.3 | 2026-01-09 | |
| 7 | 94.8% | 46.9 | 2026-01-09 | |
| 8 | 94.6% | 53.0 | 2026-01-09 | |
| 9 | 94.6% | 49.6 | 2026-01-09 | |
| 10 | 94.2% | 53.7 | 2026-01-09 | |
| 11 | 94.2% | 51.0 | 2026-01-09 | |
| 12 | 94.2% | 50.8 | 2026-01-09 | |
| 13 | 94.2% | 44.3 | 2026-01-09 | |
| 14 | 94.0% | 52.1 | 2026-01-09 | |
| 15 | 93.8% | 54.0 | 2026-01-09 | |
| 16 | 93.8% | 48.2 | 2026-01-09 | |
| 17 | 93.0% | 51.9 | 2026-01-09 | |
| 18 | 92.2% | 48.7 | 2026-01-09 | |
| 19 | 91.8% | 48.4 | 2026-01-09 | |
| 20 | 91.8% | 47.7 | 2026-01-09 | |
| 21 | 91.6% | 52.5 | 2026-01-09 | |
| 22 | 91.6% | 51.7 | 2026-01-09 | |
| 23 | 91.4% | – | 2026-01-09 | |
| 24 | 90.4% | 52.9 | 2026-01-09 | |
| 25 | 90.4% | 44.5 | 2026-01-09 | |
| 26 | 90.3% | 49.7 | 2026-01-09 | |
| 27 | 89.8% | 51.0 | 2026-01-09 | |
| 28 | 89.0% | 55.1 | 2026-01-09 | |
| 29 | 89.0% | 43.7 | 2026-01-09 | |
| 30 | 88.0% | 46.3 | 2026-01-09 | |
| 31 | 87.2% | 46.6 | 2026-01-09 | |
| 32 | 87.0% | 45.5 | 2026-01-09 | |
| 33 | 85.2% | 39.4 | 2026-01-09 | |
| 34 | 84.6% | 45.2 | 2026-01-09 | |
| 35 | 82.8% | 44.9 | 2026-01-09 | |
| 36 | 80.4% | 44.8 | 2026-01-09 | |
| 37 | 80.2% | 32.2 | 2026-01-09 | |
| 38 | 79.2% | 36.4 | 2026-01-09 | |
| 39 | 78.8% | 40.8 | 2026-01-09 | |
| 40 | 78.4% | 42.2 | 2026-01-09 | |
| 41 | 76.8% | 50.2 | 2026-01-09 | |
| 42 | 76.2% | 42.3 | 2026-01-09 | |
| 43 | 75.2% | 41.6 | 2026-01-09 | |
| 44 | 74.4% | 41.0 | 2026-01-09 | |
| 45 | 73.4% | 39.2 | 2026-01-09 | |
| 46 | 72.6% | 33.6 | 2026-01-09 | |
| 47 | 72.4% | 46.6 | 2026-01-09 | |
| 48 | 71.4% | – | 2026-01-09 | |
| 49 | 71.2% | – | 2026-01-09 | |
| 50 | 70.6% | 32.0 | 2026-01-09 | |
| 51 | 70.2% | 50.6 | 2026-01-09 | |
| 52 | 68.4% | 34.6 | 2026-01-09 | |
| 53 | 65.0% | – | 2026-01-09 | |
| 54 | 64.2% | 40.3 | 2026-01-09 | |
| 55 | 54.8% | 26.0 | 2026-01-09 | |
| 56 | 44.4% | – | 2026-01-09 | |
| 57 | 25.4% | 21.8 | 2026-01-09 |
Cite as: BenchLeader, “MATH 500 leaderboard”, https://www.benchleader.com/benchmarks/vals_math500, data as of 19 Sept 2026.
MATH 500: questions
- What does MATH 500 measure?
- MATH 500 is a 500-problem subset of the MATH competition dataset spanning algebra, counting, geometry and number theory, each with a short exact answer. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads MATH 500?
- Gemini 3 Pro leads MATH 500 with 96.4% as of 19 Sept 2026, ahead of Grok 4 at 96.2%.
- How many models have MATH 500 results?
- 57 model configurations have a MATH 500 result on BenchLeader, all taken from Vals AI.
- Who runs MATH 500 and how often is it updated?
- MATH 500 is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does MATH 500 count toward the BenchLeader Index?
- No. MATH 500 is shown for reference but left out of the composite index.