BenchLeader

BTF-3

Forecasting tournament pooled score.

As of 19 Sept 2026, Claude Sonnet 5 leads BTF-3 on BenchLeader with 15.4%, ahead of GPT-5.5 at 14.3%, across 8 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Reasoning
Index weight
Reference only
Models
8
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Binary and numeric forecasting questions scored with Brier and ranked probability scores, pooled.

How it is scored

Pooled score, as published.

What to keep in mind

Very few models; scores depend on question resolution.

8 of 8
#
1Claude Sonnet 5xhighAnthropic15.4%59.82026-06-30
2GPT-5.5highOpenAI14.3%67.02026-04-23
3Claude Opus 4.8highAnthropic14.0%62.22026-05-28
4GPT-5.6 SolmaxOpenAI13.7%68.82026-07-09
5GPT-5.6 SolhighOpenAI13.5%68.02026-07-09
6Claude Fable 5highAnthropic13.0%64.02026-06-09
7Claude Opus 4.8xhighAnthropic13.0%2026-05-28
8Claude Opus 5xhighAnthropic11.8%70.22026-07-24

Cite as: BenchLeader, “BTF-3 leaderboard”, https://www.benchleader.com/benchmarks/btf3, data as of 19 Sept 2026.

BTF-3: questions

What does BTF-3 measure?
Binary and numeric forecasting questions scored with Brier and ranked probability scores, pooled. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads BTF-3?
Claude Sonnet 5 leads BTF-3 with 15.4% as of 19 Sept 2026, ahead of GPT-5.5 at 14.3%.
How many models have BTF-3 results?
8 model configurations have a BTF-3 result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs BTF-3 and how often is it updated?
BTF-3 is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does BTF-3 count toward the BenchLeader Index?
No. BTF-3 is shown for reference but left out of the composite index.