GPQA Diamond (HELM)
Stanford CRFM's own GPQA Diamond run. Shown for reference; Epoch's run is in the index.
As of 12 Sept 2026, Gemini 3 Pro leads GPQA Diamond (HELM) on BenchLeader with 80.3%, ahead of GPT-5 at 79.1%, across 68 model configurations with a published result.
- Published by
- HELM Capabilities
- Category
- Reasoning
- Index weight
- Reference only
- Models
- 68
- Data as of
- 12 Sept 2026
Stanford Center for Research on Foundation Models (crfm.stanford.edu), Apache-2.0 results.
What the test looks like
GPQA Diamond is a set of 198 multiple-choice questions in biology, physics and chemistry written by PhD-level domain experts. The questions are designed to be "Google-proof": a skilled non-expert with unlimited web access still answers only about a third correctly, while experts in the field score around two-thirds. The Diamond subset keeps only the questions that experts answered correctly and non-experts got wrong.
How it is scored
Accuracy, the share of questions answered correctly, with four options per question so random guessing scores 25%. Epoch AI runs every model with the same prompt and reports the mean over several attempts.
What to keep in mind
Stanford CRFM's own run of GPQA Diamond in HELM's fixed harness. Shown for reference; Epoch AI's run is counted in the index. See also the GPQA Diamond board.
- 1Gemini 3 Pro80.3%
- 2GPT-579.1%
- 3GPT-5 mini75.6%
- 4o375.3%
- 5Gemini 2.5 Pro74.9%
- 6o4-mini73.5%
- 7Grok 472.6%
- 8Qwen3 235B A22B 250772.6%
- 9Claude Opus 470.9%
- 10Claude Sonnet 470.6%
- 11Claude Sonnet 4.568.6%
- 12gpt-oss-120b68.4%
- 13GPT-5 nano67.9%
- 14Grok 3 mini67.5%
- 15Claude Opus 466.6%
| # | |||||
|---|---|---|---|---|---|
| 1 | 80.3% | – | 62.0 | – | |
| 2 | 79.1% | – | 60.8 | – | |
| 3 | 75.6% | – | 57.9 | – | |
| 4 | 75.3% | – | 61.6 | – | |
| 5 | 74.9% | – | 54.5 | – | |
| 6 | 73.5% | – | 57.2 | – | |
| 7 | 72.6% | 0709 | 58.6 | – | |
| 8 | 72.6% | Qwen3 235B A22B Instruct 2507 FP8 | 54.1 | – | |
| 9 | 70.9% | – | 56.1 | – | |
| 10 | 70.6% | – | 55.1 | – | |
| 11 | 68.6% | 20250929 | 53.7 | – | |
| 12 | 68.4% | – | 48.8 | – | |
| 13 | 67.9% | – | 52.5 | – | |
| 14 | 67.5% | – | 51.8 | – | |
| 15 | 66.6% | 20250514 | 52.2 | – | |
| 16 | 66.6% | – | 49.9 | – | |
| 17 | 66.1% | Palmyra X5 | – | – | |
| 18 | 65.9% | – | 48.8 | – | |
| 19 | 65.2% | Kimi K2 Instruct | 51.2 | – | |
| 20 | 65.0% | – | 51.0 | – | |
| 21 | 65.0% | Llama 4 Maverick (17Bx128E) Instruct FP8 | 44.2 | – | |
| 22 | 64.3% | 20250514 | 50.5 | – | |
| 23 | 63.0% | – | 51.2 | – | |
| 24 | 62.3% | Qwen3 235B A22B FP8 Throughput | 44.8 | – | |
| 25 | 61.4% | – | 46.3 | – | |
| 26 | 60.8% | 20250219 | 50.2 | – | |
| 27 | 60.5% | 20251001 | 48.6 | – | |
| 28 | 59.4% | – | 47.0 | – | |
| 29 | 59.4% | – | 46.0 | – | |
| 30 | 56.5% | 20241022 | 46.3 | – | |
| 31 | 55.6% | Gemini 2.0 Flash | 43.5 | – | |
| 32 | 53.8% | DeepSeek v3 | 45.2 | – | |
| 33 | 53.4% | Gemini 1.5 Pro (002) | 44.8 | – | |
| 34 | 52.2% | Llama 3.1 Instruct Turbo (405B) | 43.1 | – | |
| 35 | 52.0% | – | 43.8 | – | |
| 36 | 51.8% | – | 45.8 | – | |
| 37 | 50.7% | Llama 4 Scout (17Bx16E) Instruct | 40.4 | – | |
| 38 | 50.7% | – | 39.1 | – | |
| 39 | 50.0% | – | 47.7 | – | |
| 40 | 44.6% | – | 41.1 | – | |
| 41 | 44.2% | – | 59.4 | – | |
| 42 | 43.7% | Gemini 1.5 Flash (002) | 40.8 | – | |
| 43 | 43.5% | 2411 | 40.8 | – | |
| 44 | 42.6% | Qwen2.5 Instruct Turbo (72B) | 42.7 | – | |
| 45 | 42.6% | Llama 3.1 Instruct Turbo (70B) | 41.1 | – | |
| 46 | 42.2% | – | – | – | |
| 47 | 39.7% | – | 39.4 | – | |
| 48 | 39.5% | – | – | – | |
| 49 | 39.2% | 2503 | 40.3 | – | |
| 50 | 39.0% | – | 50.0 | – | |
| 51 | 38.3% | – | 39.8 | – | |
| 52 | 38.3% | IBM Granite 4.0 Small | – | – | |
| 53 | 36.8% | – | 36.6 | – | |
| 54 | 36.8% | – | – | – | |
| 55 | 36.3% | 20241022 | 40.0 | – | |
| 56 | 34.1% | Qwen2.5 Instruct Turbo (7B) | 40.1 | – | |
| 57 | 33.4% | Mixtral Instruct (8x22B) | 37.6 | – | |
| 58 | 32.5% | IBM Granite 3.3 8B Instruct | – | – | |
| 59 | 31.6% | OLMo 2 13B Instruct November 2024 | – | – | |
| 60 | 30.9% | – | 43.8 | – |
Cite as: BenchLeader, “GPQA Diamond (HELM) leaderboard”, https://www.benchleader.com/benchmarks/helm_gpqa, data as of 12 Sept 2026.
GPQA Diamond (HELM): questions
- What does GPQA Diamond (HELM) measure?
- GPQA Diamond is a set of 198 multiple-choice questions in biology, physics and chemistry written by PhD-level domain experts. The questions are designed to be "Google-proof": a skilled non-expert with unlimited web access still answers only about a third correctly, while experts in the field score around two-thirds. The Diamond subset keeps only the questions that experts answered correctly and non-experts got wrong. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads GPQA Diamond (HELM)?
- Gemini 3 Pro leads GPQA Diamond (HELM) with 80.3% as of 12 Sept 2026, ahead of GPT-5 at 79.1%.
- How many models have GPQA Diamond (HELM) results?
- 68 model configurations have a GPQA Diamond (HELM) result on BenchLeader, all taken from HELM Capabilities.
- Who runs GPQA Diamond (HELM) and how often is it updated?
- GPQA Diamond (HELM) is published by HELM Capabilities. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does GPQA Diamond (HELM) count toward the BenchLeader Index?
- No. GPQA Diamond (HELM) is shown for reference but left out of the composite index.