BenchLeader

GPQA Diamond (HELM)

Stanford CRFM's own GPQA Diamond run. Shown for reference; Epoch's run is in the index.

As of 12 Sept 2026, Gemini 3 Pro leads GPQA Diamond (HELM) on BenchLeader with 80.3%, ahead of GPT-5 at 79.1%, across 68 model configurations with a published result.

Published by
HELM Capabilities
Category
Reasoning
Index weight
Reference only
Models
68
Data as of
12 Sept 2026

Stanford Center for Research on Foundation Models (crfm.stanford.edu), Apache-2.0 results.

What the test looks like

GPQA Diamond is a set of 198 multiple-choice questions in biology, physics and chemistry written by PhD-level domain experts. The questions are designed to be "Google-proof": a skilled non-expert with unlimited web access still answers only about a third correctly, while experts in the field score around two-thirds. The Diamond subset keeps only the questions that experts answered correctly and non-experts got wrong.

How it is scored

Accuracy, the share of questions answered correctly, with four options per question so random guessing scores 25%. Epoch AI runs every model with the same prompt and reports the mean over several attempts.

What to keep in mind

Stanford CRFM's own run of GPQA Diamond in HELM's fixed harness. Shown for reference; Epoch AI's run is counted in the index. See also the GPQA Diamond board.

68 of 68
#
1Gemini 3 ProGoogle80.3%62.0
2GPT-5OpenAI79.1%60.8
3GPT-5 miniOpenAI75.6%57.9
4o3OpenAI75.3%61.6
5Gemini 2.5 ProGoogle74.9%54.5
6o4-miniOpenAI73.5%57.2
7Grok 4xAI72.6%070958.6
8Qwen3 235B A22B 2507Alibabaopen72.6%Qwen3 235B A22B Instruct 2507 FP854.1
9Claude Opus 4Anthropic70.9%56.1
10Claude Sonnet 4Anthropic70.6%55.1
11Claude Sonnet 4.5Anthropic68.6%2025092953.7
12gpt-oss-120bOpenAIopen68.4%48.8
13GPT-5 nanoOpenAI67.9%52.5
14Grok 3 minixAI67.5%51.8
15Claude Opus 4Anthropic66.6%2025051452.2
16DeepSeek R1 0528DeepSeekopen66.6%49.9
17Palmyra X5Writer66.1%Palmyra X5
18GPT-4.1OpenAI65.9%48.8
19Kimi K2Moonshot AIopen65.2%Kimi K2 Instruct51.2
20Grok 3xAI65.0%51.0
21Llama 4 MaverickMetaopen65.0%Llama 4 Maverick (17Bx128E) Instruct FP844.2
22Claude Sonnet 4Anthropic64.3%2025051450.5
23Qwen3-Next 80B A3BAlibabaopen63.0%51.2
24Qwen3 235B A22BAlibabaopen62.3%Qwen3 235B A22B FP8 Throughput44.8
25GPT-4.1 miniOpenAI61.4%46.3
26Claude 3.7 SonnetAnthropic60.8%2025021950.2
27Claude Haiku 4.5Anthropic60.5%2025100148.6
28GLM-4.5-AirZhipu AIopen59.4%47.0
29gpt-oss-20bOpenAIopen59.4%46.0
30Claude 3.5 SonnetAnthropic56.5%2024102246.3
31Gemini 2.0 FlashGoogle55.6%Gemini 2.0 Flash43.5
32DeepSeek V3DeepSeekopen53.8%DeepSeek v345.2
33Gemini 1.5 Pro 002Google53.4%Gemini 1.5 Pro (002)44.8
34Llama 3.1 405BMetaopen52.2%Llama 3.1 Instruct Turbo (405B)43.1
35GPT-4oOpenAI52.0%43.8
36Nova PremierAmazon51.8%45.8
37Llama 4 ScoutMetaopen50.7%Llama 4 Scout (17Bx16E) Instruct40.4
38GPT-4.1 nanoOpenAI50.7%39.1
39Gemini 2.0 Flash-LiteGoogle50.0%47.7
40Nova ProAmazon44.6%41.1
41GPT-5.1OpenAI44.2%59.4
42Gemini 1.5 Flash 002Google43.7%Gemini 1.5 Flash (002)40.8
43Mistral Large 2Mistral AIopen43.5%241140.8
44Qwen2.5 72BAlibabaopen42.6%Qwen2.5 Instruct Turbo (72B)42.7
45Llama 3.1 70BMetaopen42.6%Llama 3.1 Instruct Turbo (70B)41.1
46Palmyra FinWriter42.2%
47Amazon Nova LiteAmazon39.7%39.4
48Palmyra X 004Writer39.5%
49Mistral Small 3.1Mistral AIopen39.2%250340.3
50Gemini 2.5 FlashGoogle39.0%50.0
51Nova MicroAmazon38.3%39.8
52Granite 4.0 SmallIBM38.3%IBM Granite 4.0 Small
53GPT-4o miniOpenAI36.8%36.6
54Palmyra MedWriter36.8%
55Claude 3.5 HaikuAnthropic36.3%2024102240.0
56Qwen2.5 7B InstructAlibabaopen34.1%Qwen2.5 Instruct Turbo (7B)40.1
57Mixtral 8x22B InstructMistral AIopen33.4%Mixtral Instruct (8x22B)37.6
58Granite 3.3 8BIBM32.5%IBM Granite 3.3 8B Instruct
59Olmo 2 13B November 2024Ai231.6%OLMo 2 13B Instruct November 2024
60Gemini 2.5 Flash-LiteGoogle30.9%43.8

Cite as: BenchLeader, “GPQA Diamond (HELM) leaderboard”, https://www.benchleader.com/benchmarks/helm_gpqa, data as of 12 Sept 2026.

GPQA Diamond (HELM): questions

What does GPQA Diamond (HELM) measure?
GPQA Diamond is a set of 198 multiple-choice questions in biology, physics and chemistry written by PhD-level domain experts. The questions are designed to be "Google-proof": a skilled non-expert with unlimited web access still answers only about a third correctly, while experts in the field score around two-thirds. The Diamond subset keeps only the questions that experts answered correctly and non-experts got wrong. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads GPQA Diamond (HELM)?
Gemini 3 Pro leads GPQA Diamond (HELM) with 80.3% as of 12 Sept 2026, ahead of GPT-5 at 79.1%.
How many models have GPQA Diamond (HELM) results?
68 model configurations have a GPQA Diamond (HELM) result on BenchLeader, all taken from HELM Capabilities.
Who runs GPQA Diamond (HELM) and how often is it updated?
GPQA Diamond (HELM) is published by HELM Capabilities. BenchLeader re-reads the published results every morning and records the date each result was published.
Does GPQA Diamond (HELM) count toward the BenchLeader Index?
No. GPQA Diamond (HELM) is shown for reference but left out of the composite index.