Omni-MATH (HELM)
Olympiad-level maths problems. Run by Stanford CRFM's HELM.
As of 12 Sept 2026, GPT-5 mini leads Omni-MATH (HELM) on BenchLeader with 72.2%, ahead of o4-mini at 72.0%, across 68 model configurations with a published result.
- Published by
- HELM Capabilities
- Category
- Maths
- Index weight
- 1.0
- Models
- 68
- Data as of
- 12 Sept 2026
Stanford Center for Research on Foundation Models (crfm.stanford.edu), Apache-2.0 results.
What the test looks like
Omni-MATH is a set of olympiad-level mathematics problems across many topics and difficulty tiers, with answers checked against the reference.
How it is scored
Accuracy on the problems, run by Stanford CRFM's HELM with the same prompt for every model.
What to keep in mind
Answer checking for olympiad problems is imperfect, and HELM runs models without extended reasoning budgets unless the listing says so, so scores sit below what labs report with heavy test-time compute.
- 1GPT-5 mini72.2%
- 2o4-mini72.0%
- 3Qwen3 235B A22B 250771.8%
- 4o371.4%
- 5gpt-oss-120b68.8%
- 6Kimi K265.4%
- 7GPT-564.7%
- 8Claude Opus 461.6%
- 9Grok 460.3%
- 10Claude Sonnet 460.2%
- 11gpt-oss-20b56.5%
- 12Claude Haiku 4.556.1%
- 13Gemini 3 Pro55.6%
- 14Claude Sonnet 4.555.3%
- 15Qwen3 235B A22B54.8%
| # | |||||
|---|---|---|---|---|---|
| 1 | 72.2% | – | 57.9 | – | |
| 2 | 72.0% | – | 57.2 | – | |
| 3 | 71.8% | Qwen3 235B A22B Instruct 2507 FP8 | 54.1 | – | |
| 4 | 71.4% | – | 61.6 | – | |
| 5 | 68.8% | – | 48.8 | – | |
| 6 | 65.4% | Kimi K2 Instruct | 51.2 | – | |
| 7 | 64.7% | – | 60.8 | – | |
| 8 | 61.6% | – | 56.1 | – | |
| 9 | 60.3% | 0709 | 58.6 | – | |
| 10 | 60.2% | – | 55.1 | – | |
| 11 | 56.5% | – | 46.0 | – | |
| 12 | 56.1% | 20251001 | 48.6 | – | |
| 13 | 55.6% | – | 62.0 | – | |
| 14 | 55.3% | 20250929 | 53.7 | – | |
| 15 | 54.8% | Qwen3 235B A22B FP8 Throughput | 44.8 | – | |
| 16 | 54.7% | – | 52.5 | – | |
| 17 | 51.3% | 20250514 | 50.5 | – | |
| 18 | 51.1% | 20250514 | 52.2 | – | |
| 19 | 49.1% | – | 46.3 | – | |
| 20 | 48.0% | – | 43.8 | – | |
| 21 | 47.1% | – | 48.8 | – | |
| 22 | 46.7% | – | 51.2 | – | |
| 23 | 46.4% | – | 59.4 | – | |
| 24 | 46.4% | – | 51.0 | – | |
| 25 | 45.9% | Gemini 2.0 Flash | 43.5 | – | |
| 26 | 42.4% | – | 49.9 | – | |
| 27 | 42.2% | Llama 4 Maverick (17Bx128E) Instruct FP8 | 44.2 | – | |
| 28 | 41.6% | – | 54.5 | – | |
| 29 | 41.5% | Palmyra X5 | – | – | |
| 30 | 40.3% | DeepSeek v3 | 45.2 | – | |
| 31 | 39.1% | – | 47.0 | – | |
| 32 | 38.4% | – | 50.0 | – | |
| 33 | 37.4% | – | 47.7 | – | |
| 34 | 37.3% | Llama 4 Scout (17Bx16E) Instruct | 40.4 | – | |
| 35 | 36.7% | – | 39.1 | – | |
| 36 | 36.4% | Gemini 1.5 Pro (002) | 44.8 | – | |
| 37 | 35.0% | – | 45.8 | – | |
| 38 | 33.0% | 20250219 | 50.2 | – | |
| 39 | 33.0% | Qwen2.5 Instruct Turbo (72B) | 42.7 | – | |
| 40 | 32.0% | – | – | – | |
| 41 | 31.8% | – | 51.8 | – | |
| 42 | 30.5% | Gemini 1.5 Flash (002) | 40.8 | – | |
| 43 | 29.6% | IBM Granite 4.0 Small | – | – | |
| 44 | 29.5% | – | – | – | |
| 45 | 29.4% | Qwen2.5 Instruct Turbo (7B) | 40.1 | – | |
| 46 | 29.3% | – | 43.8 | – | |
| 47 | 28.1% | 2411 | 40.8 | – | |
| 48 | 28.0% | – | 36.6 | – | |
| 49 | 27.6% | 20241022 | 46.3 | – | |
| 50 | 24.9% | Llama 3.1 Instruct Turbo (405B) | 43.1 | – | |
| 51 | 24.8% | 2503 | 40.3 | – | |
| 52 | 24.2% | – | 41.1 | – | |
| 53 | 23.3% | – | 39.4 | – | |
| 54 | 22.4% | 20241022 | 40.0 | – | |
| 55 | 21.4% | – | 39.8 | – | |
| 56 | 21.0% | Llama 3.1 Instruct Turbo (70B) | 41.1 | – | |
| 57 | 20.9% | IBM Granite 4.0 Micro | 35.9 | – | |
| 58 | 17.6% | IBM Granite 3.3 8B Instruct | – | – | |
| 59 | 16.3% | Mixtral Instruct (8x22B) | 37.6 | – | |
| 60 | 16.1% | OLMo 2 32B Instruct March 2025 | – | – |
Cite as: BenchLeader, “Omni-MATH (HELM) leaderboard”, https://www.benchleader.com/benchmarks/helm_omni_math, data as of 12 Sept 2026.
Omni-MATH (HELM): questions
- What does Omni-MATH (HELM) measure?
- Omni-MATH is a set of olympiad-level mathematics problems across many topics and difficulty tiers, with answers checked against the reference. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads Omni-MATH (HELM)?
- GPT-5 mini leads Omni-MATH (HELM) with 72.2% as of 12 Sept 2026, ahead of o4-mini at 72.0%.
- How many models have Omni-MATH (HELM) results?
- 68 model configurations have a Omni-MATH (HELM) result on BenchLeader, all taken from HELM Capabilities.
- Who runs Omni-MATH (HELM) and how often is it updated?
- Omni-MATH (HELM) is published by HELM Capabilities. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does Omni-MATH (HELM) count toward the BenchLeader Index?
- Yes. Omni-MATH (HELM) contributes to the maths category of the BenchLeader Index, normalised so that 50 is the average of the evaluated models.