BenchLeader

EBR-bench

Evidence-based reasoning over research literature. Run by Epoch AI.

As of 19 Sept 2026, GPT-6 Astra leads EBR-bench on BenchLeader with 76.2%, ahead of Claude Fable 5.1 at 57.1%, across 21 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Knowledge
Index weight
Reference only
Models
21
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Questions that require reading and weighing research evidence to reach a supported conclusion.

How it is scored

Percent correct, run by Epoch AI.

What to keep in mind

Small set; new.

21 of 21
#
1GPT-6 AstramaxOpenAI76.2%71.82026-09-03
2Claude Fable 5.1maxAnthropic57.1%69.72026-09-01
3Claude Opus 5maxAnthropic45.7%69.92026-07-24
4GPT-5.6 SolmaxOpenAI44.8%68.82026-07-09
5Claude Fable 5maxAnthropic39.5%65.82026-06-09
6GPT-5.5xhighOpenAI34.3%67.62026-04-23
7Grok 4.6xhighxAI30.5%64.62026-08-12
8Claude Opus 4.8maxAnthropic28.6%64.22026-05-28
9GPT-5.4xhighOpenAI25.4%65.42026-03-05
10GPT-5.2xhighOpenAI23.0%62.02025-12-11
11Claude Opus 4.7maxAnthropic19.1%64.32026-04-16
12Gemini 3.1 ProGoogle14.3%63.92026-02-19
13Claude Opus 4.5Anthropic14.3%58.22025-11-24
14Claude Opus 4.6maxAnthropic12.7%59.22026-02-05
15GPT-5highOpenAI12.7%58.62025-08-07
16GLM-5.2maxZhipu AIopen ↗9.5%63.92026-06-16
17Qwen3 7maxAlibaba9.5%61.02026-05-19
18Claude Opus 4.1Anthropic7.9%51.92025-08-05
19Gemini 3.5 FlashhighGoogle4.8%63.62026-05-19
20Kimi K2.6Moonshot AIopen ↗2.4%60.52026-04-20
21Claude Sonnet 4.5Anthropic2.4%54.32025-09-29

Cite as: BenchLeader, “EBR-bench leaderboard”, https://www.benchleader.com/benchmarks/ebr_bench, data as of 19 Sept 2026.

EBR-bench: questions

What does EBR-bench measure?
Questions that require reading and weighing research evidence to reach a supported conclusion. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads EBR-bench?
GPT-6 Astra leads EBR-bench with 76.2% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 57.1%.
How many models have EBR-bench results?
21 model configurations have a EBR-bench result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs EBR-bench and how often is it updated?
EBR-bench is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does EBR-bench count toward the BenchLeader Index?
No. EBR-bench is shown for reference but left out of the composite index.