BenchLeader

HELM Capabilities mean

Mean of HELM's five capability scenarios. Shown for reference.

As of 12 Sept 2026, GPT-5 mini leads HELM Capabilities mean on BenchLeader with 81.9%, ahead of o4-mini at 81.2%, across 68 model configurations with a published result.

Published by
HELM Capabilities
Category
Composite
Index weight
Reference only
Models
68
Data as of
12 Sept 2026

Stanford Center for Research on Foundation Models (crfm.stanford.edu), Apache-2.0 results.

What the test looks like

The mean of HELM's five capability scenarios: MMLU-Pro, GPQA, IFEval, WildBench and Omni-MATH.

How it is scored

Unweighted mean of the five scores, on a 0 to 100 scale. Shown for reference only.

What to keep in mind

Five scenarios is a narrow base for a composite, and HELM updates roughly monthly, so new models appear with a lag.

68 of 68
#
1GPT-5 miniOpenAI81.9%57.9
2o4-miniOpenAI81.2%57.2
3o3OpenAI81.1%61.6
4GPT-5OpenAI80.7%60.8
5Gemini 3 ProGoogle79.9%62.0
6Qwen3 235B A22B 2507Alibabaopen79.8%Qwen3 235B A22B Instruct 2507 FP854.1
7Grok 4xAI78.5%070958.6
8Claude Opus 4Anthropic78.0%56.1
9gpt-oss-120bOpenAIopen77.0%48.8
10Kimi K2Moonshot AIopen76.8%Kimi K2 Instruct51.2
11Claude Sonnet 4Anthropic76.6%55.1
12Claude Sonnet 4.5Anthropic76.2%2025092953.7
13Claude Opus 4Anthropic75.7%2025051452.2
14GPT-5 nanoOpenAI74.8%52.5
15Gemini 2.5 ProGoogle74.5%54.5
16Claude Sonnet 4Anthropic73.3%2025051450.5
17Grok 3xAI72.7%51.0
18GPT-4.1OpenAI72.7%48.8
19GPT-4.1 miniOpenAI72.6%46.3
20Qwen3 235B A22BAlibabaopen72.6%Qwen3 235B A22B FP8 Throughput44.8
21Llama 4 MaverickMetaopen71.8%Llama 4 Maverick (17Bx128E) Instruct FP844.2
22Claude Haiku 4.5Anthropic71.7%2025100148.6
23Qwen3-Next 80B A3BAlibabaopen70.0%51.2
24DeepSeek R1 0528DeepSeekopen69.9%49.9
25Palmyra X5Writer69.6%Palmyra X5
26Grok 3 minixAI67.9%51.8
27Gemini 2.0 FlashGoogle67.9%Gemini 2.0 Flash43.5
28Claude 3.7 SonnetAnthropic67.4%2025021950.2
29gpt-oss-20bOpenAIopen67.4%46.0
30GLM-4.5-AirZhipu AIopen67.0%47.0
31DeepSeek V3DeepSeekopen66.5%DeepSeek v345.2
32Gemini 1.5 Pro 002Google65.7%Gemini 1.5 Pro (002)44.8
33GPT-5.1OpenAI65.6%59.4
34Claude 3.5 SonnetAnthropic65.3%2024102246.3
35Llama 4 ScoutMetaopen64.4%Llama 4 Scout (17Bx16E) Instruct40.4
36Gemini 2.0 Flash-LiteGoogle64.2%47.7
37Nova PremierAmazon63.7%45.8
38GPT-4oOpenAI63.4%43.8
39Gemini 2.5 FlashGoogle62.6%50.0
40Llama 3.1 405BMetaopen61.8%Llama 3.1 Instruct Turbo (405B)43.1
41GPT-4.1 nanoOpenAI61.6%39.1
42Gemini 1.5 Flash 002Google60.9%Gemini 1.5 Flash (002)40.8
43Palmyra X 004Writer60.9%
44Qwen2.5 72BAlibabaopen59.9%Qwen2.5 Instruct Turbo (72B)42.7
45Mistral Large 2Mistral AIopen59.8%241140.8
46Gemini 2.5 Flash-LiteGoogle59.1%43.8
47Nova ProAmazon59.1%41.1
48Palmyra FinWriter57.7%
49Granite 4.0 SmallIBM57.5%IBM Granite 4.0 Small
50Llama 3.1 70BMetaopen57.4%Llama 3.1 Instruct Turbo (70B)41.1
51GPT-4o miniOpenAI56.5%36.6
52Mistral Small 3.1Mistral AIopen55.8%250340.3
53Amazon Nova LiteAmazon55.1%39.4
54Claude 3.5 HaikuAnthropic54.9%2024102240.0
55Qwen2.5 7B InstructAlibabaopen52.9%Qwen2.5 Instruct Turbo (7B)40.1
56Nova MicroAmazon52.2%39.8
57granite-4.0-microIBMopen48.6%IBM Granite 4.0 Micro35.9
58Mixtral 8x22B InstructMistral AIopen47.8%Mixtral Instruct (8x22B)37.6
59Palmyra MedWriter47.6%
60Olmo 2 32B March 2025Ai247.5%OLMo 2 32B Instruct March 2025

Cite as: BenchLeader, “HELM Capabilities mean leaderboard”, https://www.benchleader.com/benchmarks/helm_mean, data as of 12 Sept 2026.

HELM Capabilities mean: questions

What does HELM Capabilities mean measure?
The mean of HELM's five capability scenarios: MMLU-Pro, GPQA, IFEval, WildBench and Omni-MATH. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads HELM Capabilities mean?
GPT-5 mini leads HELM Capabilities mean with 81.9% as of 12 Sept 2026, ahead of o4-mini at 81.2%.
How many models have HELM Capabilities mean results?
68 model configurations have a HELM Capabilities mean result on BenchLeader, all taken from HELM Capabilities.
Who runs HELM Capabilities mean and how often is it updated?
HELM Capabilities mean is published by HELM Capabilities. BenchLeader re-reads the published results every morning and records the date each result was published.
Does HELM Capabilities mean count toward the BenchLeader Index?
No. HELM Capabilities mean is shown for reference but left out of the composite index.