EBR-bench
Evidence-based reasoning over research literature. Run by Epoch AI.
As of 19 Sept 2026, GPT-6 Astra leads EBR-bench on BenchLeader with 76.2%, ahead of Claude Fable 5.1 at 57.1%, across 21 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Knowledge
- Index weight
- Reference only
- Models
- 21
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Questions that require reading and weighing research evidence to reach a supported conclusion.
How it is scored
Percent correct, run by Epoch AI.
What to keep in mind
Small set; new.
- 1GPT-6 Astra (max)76.2%
- 2Claude Fable 5.1 (max)57.1%
- 3Claude Opus 5 (max)45.7%
- 4GPT-5.6 Sol (max)44.8%
- 5Claude Fable 5 (max)39.5%
- 6GPT-5.5 (xhigh)34.3%
- 7Grok 4.6 (xhigh)30.5%
- 8Claude Opus 4.8 (max)28.6%
- 9GPT-5.4 (xhigh)25.4%
- 10GPT-5.2 (xhigh)23.0%
- 11Claude Opus 4.7 (max)19.1%
- 12Gemini 3.1 Pro14.3%
- 13Claude Opus 4.514.3%
- 14Claude Opus 4.6 (max)12.7%
- 15GPT-5 (high)12.7%
21 of 21
| # | ||||
|---|---|---|---|---|
| 1 | 76.2% | 71.8 | 2026-09-03 | |
| 2 | 57.1% | 69.7 | 2026-09-01 | |
| 3 | 45.7% | 69.9 | 2026-07-24 | |
| 4 | 44.8% | 68.8 | 2026-07-09 | |
| 5 | 39.5% | 65.8 | 2026-06-09 | |
| 6 | 34.3% | 67.6 | 2026-04-23 | |
| 7 | 30.5% | 64.6 | 2026-08-12 | |
| 8 | 28.6% | 64.2 | 2026-05-28 | |
| 9 | 25.4% | 65.4 | 2026-03-05 | |
| 10 | 23.0% | 62.0 | 2025-12-11 | |
| 11 | 19.1% | 64.3 | 2026-04-16 | |
| 12 | 14.3% | 63.9 | 2026-02-19 | |
| 13 | 14.3% | 58.2 | 2025-11-24 | |
| 14 | 12.7% | 59.2 | 2026-02-05 | |
| 15 | 12.7% | 58.6 | 2025-08-07 | |
| 16 | 9.5% | 63.9 | 2026-06-16 | |
| 17 | 9.5% | 61.0 | 2026-05-19 | |
| 18 | 7.9% | 51.9 | 2025-08-05 | |
| 19 | 4.8% | 63.6 | 2026-05-19 | |
| 20 | 2.4% | 60.5 | 2026-04-20 | |
| 21 | 2.4% | 54.3 | 2025-09-29 |
Cite as: BenchLeader, “EBR-bench leaderboard”, https://www.benchleader.com/benchmarks/ebr_bench, data as of 19 Sept 2026.
EBR-bench: questions
- What does EBR-bench measure?
- Questions that require reading and weighing research evidence to reach a supported conclusion. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads EBR-bench?
- GPT-6 Astra leads EBR-bench with 76.2% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 57.1%.
- How many models have EBR-bench results?
- 21 model configurations have a EBR-bench result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs EBR-bench and how often is it updated?
- EBR-bench is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does EBR-bench count toward the BenchLeader Index?
- No. EBR-bench is shown for reference but left out of the composite index.