DeepResearch Bench
Long-form research reports judged for depth and accuracy.
As of 19 Sept 2026, Claude Opus 4.6 leads DeepResearch Bench on BenchLeader with 55.3%, ahead of Claude Sonnet 4.6 at 54.9%, across 40 model configurations with a published result.
- Published by
- DeepResearch Benchdata via Epoch AI Benchmarking Hub
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 40
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Research questions that require gathering sources and writing a report, judged by rubric.
How it is scored
Average rubric score, as published by DeepResearch Bench.
What to keep in mind
Rubric-graded reports; harness and search tools matter.
- 1Claude Opus 4.6 (high)55.3%
- 2Claude Sonnet 4.6 (high)54.9%
- 3Claude Opus 4.5 (high)54.8%
- 4GPT-5.5 (high)54.0%
- 5Claude Opus 4.5 (low)53.7%
- 6Claude Opus 4.6 (medium)53.2%
- 7Claude Sonnet 4.552.6%
- 8Claude Opus 4.6 (low)51.4%
- 9Claude Sonnet 4.6 (low)50.4%
- 10Claude Opus 4.8 (high)50.2%
- 11Gemini 3 Flash (low)49.8%
- 12GPT-5.5 (medium)49.6%
- 13GPT-5 (low)49.6%
- 14Claude Opus 4.8 (low)49.3%
- 15Gemini 3 Flash (minimal)49.0%
40 of 40
| # | ||||
|---|---|---|---|---|
| 1 | 55.3% | 61.3 | 2026-02-05 | |
| 2 | 54.9% | 54.9 | 2026-02-17 | |
| 3 | 54.8% | 57.7 | 2025-11-24 | |
| 4 | 54.0% | 67.0 | 2026-04-23 | |
| 5 | 53.7% | – | 2025-11-24 | |
| 6 | 53.2% | – | 2026-02-05 | |
| 7 | 52.6% | 54.3 | 2025-09-29 | |
| 8 | 51.4% | – | 2026-02-05 | |
| 9 | 50.4% | 54.5 | 2026-02-17 | |
| 10 | 50.2% | 62.2 | 2026-05-28 | |
| 11 | 49.8% | – | 2025-12-17 | |
| 12 | 49.6% | 64.6 | 2026-04-23 | |
| 13 | 49.6% | 53.5 | 2025-08-07 | |
| 14 | 49.3% | – | 2026-05-28 | |
| 15 | 49.0% | 54.9 | 2025-12-17 | |
| 16 | 48.9% | 46.0 | 2025-08-07 | |
| 17 | 48.7% | 61.2 | 2026-04-23 | |
| 18 | 48.6% | 58.5 | 2025-08-07 | |
| 19 | 48.3% | 51.9 | 2025-08-05 | |
| 20 | 48.1% | 58.6 | 2025-08-07 | |
| 21 | 47.9% | 58.7 | 2025-12-17 | |
| 22 | 47.8% | 60.7 | 2026-02-19 | |
| 23 | 47.5% | – | 2025-09-29 | |
| 24 | 47.4% | – | 2026-05-28 | |
| 25 | 47.3% | 58.5 | 2025-07-09 | |
| 26 | 46.8% | 52.9 | 2025-05-22 | |
| 27 | 46.6% | 49.7 | 2025-05-22 | |
| 28 | 46.3% | 54.7 | 2025-11-18 | |
| 29 | 45.5% | – | 2025-10-15 | |
| 30 | 45.2% | 53.7 | 2025-04-16 | |
| 31 | 44.5% | – | 2026-02-19 | |
| 32 | 43.6% | 50.2 | 2025-02-24 | |
| 33 | 42.8% | 54.5 | 2025-06-05 | |
| 34 | 42.8% | – | 2025-11-13 | |
| 35 | 41.1% | 50.3 | 2025-12-11 | |
| 36 | 37.3% | – | 2026-03-03 | |
| 37 | 36.4% | 54.5 | 2026-03-03 | |
| 38 | 36.3% | – | 2026-03-17 | |
| 39 | 35.1% | 60.8 | 2026-03-05 | |
| 40 | 35.1% | 50.2 | 2025-05-28 |
Cite as: BenchLeader, “DeepResearch Bench leaderboard”, https://www.benchleader.com/benchmarks/deepresearch_bench, data as of 19 Sept 2026.
DeepResearch Bench: questions
- What does DeepResearch Bench measure?
- Research questions that require gathering sources and writing a report, judged by rubric. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads DeepResearch Bench?
- Claude Opus 4.6 leads DeepResearch Bench with 55.3% as of 19 Sept 2026, ahead of Claude Sonnet 4.6 at 54.9%.
- How many models have DeepResearch Bench results?
- 40 model configurations have a DeepResearch Bench result on BenchLeader, all taken from DeepResearch Bench via Epoch AI Benchmarking Hub.
- Who runs DeepResearch Bench and how often is it updated?
- DeepResearch Bench is published by DeepResearch Bench. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does DeepResearch Bench count toward the BenchLeader Index?
- No. DeepResearch Bench is shown for reference but left out of the composite index.