BenchLeader

Omni-MATH (HELM)

Olympiad-level maths problems. Run by Stanford CRFM's HELM.

As of 12 Sept 2026, GPT-5 mini leads Omni-MATH (HELM) on BenchLeader with 72.2%, ahead of o4-mini at 72.0%, across 68 model configurations with a published result.

Published by
HELM Capabilities
Category
Maths
Index weight
1.0
Models
68
Data as of
12 Sept 2026

Stanford Center for Research on Foundation Models (crfm.stanford.edu), Apache-2.0 results.

What the test looks like

Omni-MATH is a set of olympiad-level mathematics problems across many topics and difficulty tiers, with answers checked against the reference.

How it is scored

Accuracy on the problems, run by Stanford CRFM's HELM with the same prompt for every model.

What to keep in mind

Answer checking for olympiad problems is imperfect, and HELM runs models without extended reasoning budgets unless the listing says so, so scores sit below what labs report with heavy test-time compute.

68 of 68
#
1GPT-5 miniOpenAI72.2%57.9
2o4-miniOpenAI72.0%57.2
3Qwen3 235B A22B 2507Alibabaopen71.8%Qwen3 235B A22B Instruct 2507 FP854.1
4o3OpenAI71.4%61.6
5gpt-oss-120bOpenAIopen68.8%48.8
6Kimi K2Moonshot AIopen65.4%Kimi K2 Instruct51.2
7GPT-5OpenAI64.7%60.8
8Claude Opus 4Anthropic61.6%56.1
9Grok 4xAI60.3%070958.6
10Claude Sonnet 4Anthropic60.2%55.1
11gpt-oss-20bOpenAIopen56.5%46.0
12Claude Haiku 4.5Anthropic56.1%2025100148.6
13Gemini 3 ProGoogle55.6%62.0
14Claude Sonnet 4.5Anthropic55.3%2025092953.7
15Qwen3 235B A22BAlibabaopen54.8%Qwen3 235B A22B FP8 Throughput44.8
16GPT-5 nanoOpenAI54.7%52.5
17Claude Sonnet 4Anthropic51.3%2025051450.5
18Claude Opus 4Anthropic51.1%2025051452.2
19GPT-4.1 miniOpenAI49.1%46.3
20Gemini 2.5 Flash-LiteGoogle48.0%43.8
21GPT-4.1OpenAI47.1%48.8
22Qwen3-Next 80B A3BAlibabaopen46.7%51.2
23GPT-5.1OpenAI46.4%59.4
24Grok 3xAI46.4%51.0
25Gemini 2.0 FlashGoogle45.9%Gemini 2.0 Flash43.5
26DeepSeek R1 0528DeepSeekopen42.4%49.9
27Llama 4 MaverickMetaopen42.2%Llama 4 Maverick (17Bx128E) Instruct FP844.2
28Gemini 2.5 ProGoogle41.6%54.5
29Palmyra X5Writer41.5%Palmyra X5
30DeepSeek V3DeepSeekopen40.3%DeepSeek v345.2
31GLM-4.5-AirZhipu AIopen39.1%47.0
32Gemini 2.5 FlashGoogle38.4%50.0
33Gemini 2.0 Flash-LiteGoogle37.4%47.7
34Llama 4 ScoutMetaopen37.3%Llama 4 Scout (17Bx16E) Instruct40.4
35GPT-4.1 nanoOpenAI36.7%39.1
36Gemini 1.5 Pro 002Google36.4%Gemini 1.5 Pro (002)44.8
37Nova PremierAmazon35.0%45.8
38Claude 3.7 SonnetAnthropic33.0%2025021950.2
39Qwen2.5 72BAlibabaopen33.0%Qwen2.5 Instruct Turbo (72B)42.7
40Palmyra X 004Writer32.0%
41Grok 3 minixAI31.8%51.8
42Gemini 1.5 Flash 002Google30.5%Gemini 1.5 Flash (002)40.8
43Granite 4.0 SmallIBM29.6%IBM Granite 4.0 Small
44Palmyra FinWriter29.5%
45Qwen2.5 7B InstructAlibabaopen29.4%Qwen2.5 Instruct Turbo (7B)40.1
46GPT-4oOpenAI29.3%43.8
47Mistral Large 2Mistral AIopen28.1%241140.8
48GPT-4o miniOpenAI28.0%36.6
49Claude 3.5 SonnetAnthropic27.6%2024102246.3
50Llama 3.1 405BMetaopen24.9%Llama 3.1 Instruct Turbo (405B)43.1
51Mistral Small 3.1Mistral AIopen24.8%250340.3
52Nova ProAmazon24.2%41.1
53Amazon Nova LiteAmazon23.3%39.4
54Claude 3.5 HaikuAnthropic22.4%2024102240.0
55Nova MicroAmazon21.4%39.8
56Llama 3.1 70BMetaopen21.0%Llama 3.1 Instruct Turbo (70B)41.1
57granite-4.0-microIBMopen20.9%IBM Granite 4.0 Micro35.9
58Granite 3.3 8BIBM17.6%IBM Granite 3.3 8B Instruct
59Mixtral 8x22B InstructMistral AIopen16.3%Mixtral Instruct (8x22B)37.6
60Olmo 2 32B March 2025Ai216.1%OLMo 2 32B Instruct March 2025

Cite as: BenchLeader, “Omni-MATH (HELM) leaderboard”, https://www.benchleader.com/benchmarks/helm_omni_math, data as of 12 Sept 2026.

Omni-MATH (HELM): questions

What does Omni-MATH (HELM) measure?
Omni-MATH is a set of olympiad-level mathematics problems across many topics and difficulty tiers, with answers checked against the reference. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads Omni-MATH (HELM)?
GPT-5 mini leads Omni-MATH (HELM) with 72.2% as of 12 Sept 2026, ahead of o4-mini at 72.0%.
How many models have Omni-MATH (HELM) results?
68 model configurations have a Omni-MATH (HELM) result on BenchLeader, all taken from HELM Capabilities.
Who runs Omni-MATH (HELM) and how often is it updated?
Omni-MATH (HELM) is published by HELM Capabilities. BenchLeader re-reads the published results every morning and records the date each result was published.
Does Omni-MATH (HELM) count toward the BenchLeader Index?
Yes. Omni-MATH (HELM) contributes to the maths category of the BenchLeader Index, normalised so that 50 is the average of the evaluated models.