BenchLeader

ForecastBench

Forecasting real future events, scored as they resolve.

As of 19 Sept 2026, o3 leads ForecastBench on BenchLeader with 62.5%, ahead of Claude Opus 4.1 at 62.0%, across 73 model configurations with a published result.

Published by
ForecastBenchdata via Epoch AI Benchmarking Hub
Category
Reasoning
Index weight
Reference only
Models
73
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Models forecast real-world questions; scores accrue as the questions resolve, compared with superforecasters.

How it is scored

Overall Brier-based score, as published by ForecastBench.

What to keep in mind

Scores change as events resolve; a lower Brier score is better and the site rescales it.

73 of 73
#
1o3OpenAI62.5%60.32025-04-16
2Claude Opus 4.1Anthropic62.0%51.92025-08-05
3Claude Sonnet 4.5Anthropic61.9%54.32025-09-29
4o4-miniOpenAI61.8%57.42025-04-16
5Claude 3.7 SonnetAnthropic61.8%50.22025-02-24
6GPT-4.5OpenAI61.7%50.22025-02-27
7GPT-4.1OpenAI61.5%49.02025-04-14
8GPT-5OpenAI61.4%57.22025-08-07
9Claude Haiku 4.5Anthropic61.4%48.12025-10-15
10Grok 4.20xAI61.4%2026-02-17
11Gemini 2.5 ProGoogle61.3%54.52025-03-31
12Gemini 3 ProGoogle61.2%61.12025-11-18
13Claude Sonnet 4.6Anthropic61.2%59.02026-02-17
14GPT-5.5OpenAI61.1%63.22026-04-23
15Claude Opus 4Anthropic61.1%52.92025-05-22
16GPT-5 miniOpenAI61.0%56.82025-08-07
17Grok 4.1thinkingxAI61.0%56.02025-11-19
18GLM-5Zhipu AIopen61.0%54.72026-02-11
19Grok 4xAI60.9%58.52025-07-09
20Grok 4.20thinkingxAI60.7%59.92026-02-17
21Claude Opus 4.5Anthropic60.7%58.22025-11-24
22Claude 3.5 SonnetAnthropic60.7%46.62024-10-22
23Gemini 2.5 FlashGoogle60.6%52.52025-04-17
24Grok 4 FastxAI60.5%54.32025-09-19
25MiniMax-M3MiniMaxopen ↗60.4%57.52026-06-01
26Grok 4.3xAI60.4%50.52026-04-17
27Kimi K2Moonshot AIopen60.2%51.02025-07-12
28Claude Sonnet 4Anthropic60.2%49.72025-05-22
29GPT-5.2OpenAI60.1%58.42025-12-11
30Claude Opus 4.6Anthropic60.0%63.72026-02-05
31DeepSeek R1DeepSeekopen60.0%48.72025-01-20
32Claude Sonnet 5Anthropic59.9%57.22026-06-30
33Llama 3.1 405BMetaopen59.9%43.52024-07-23
34Kimi K2 0905Moonshot AIopen59.8%51.32025-09-05
35Qwen3 235B A22BAlibabaopen59.7%49.62025-04-29
36Claude Opus 4.7Anthropic59.6%64.52026-04-16
37o3-miniOpenAI59.6%50.42025-01-31
38GPT-5.4OpenAI59.5%59.22026-03-05
39GPT-4 TurboOpenAI59.4%41.52024-04-09
40Claude Opus 4.8Anthropic59.3%61.92026-05-28
41Gemini 3.5 FlashGoogle59.2%51.32026-05-19
42GLM-4.5-AirZhipu AIopen59.2%47.22025-07-20
43GPT-5 nanoOpenAI59.1%55.92025-08-07
44DeepSeek V3DeepSeekopen59.1%44.82024-12-26
45Gemini 3.1 ProGoogle59.0%63.92026-02-19
46Llama 3.3 70BMetaopen58.6%41.42024-12-06
47Llama 3-8BMetaopen58.6%35.62024-04-18
48Gemini 3 FlashGoogle58.5%57.72025-12-17
49Gemini 1.5 Pro 001Google58.5%45.12024-05-14
50Claude 3 OpusAnthropic58.4%39.62024-02-29
51QwQ-32BAlibabaopen58.3%46.12024-11-28
52GPT-5.1OpenAI58.1%57.22025-11-13
53GPT 4 0613OpenAI57.8%40.22023-06-13
54GPT-4oOpenAI57.7%44.32024-05-13
55Qwen1 5 110BAlibabaopen57.7%42.72024-04-25
56Llama 4 MaverickMetaopen57.5%43.42025-04-05
57Qwen2.5 72BAlibabaopen57.5%42.72024-09-19
58Llama 4 ScoutMetaopen57.5%40.22025-04-05
59GPT-5.4 nanoOpenAI57.3%2026-03-17
60Mistral Large 2Mistral AIopen57.1%41.02024-07-24

Cite as: BenchLeader, “ForecastBench leaderboard”, https://www.benchleader.com/benchmarks/forecastbench, data as of 19 Sept 2026.

ForecastBench: questions

What does ForecastBench measure?
Models forecast real-world questions; scores accrue as the questions resolve, compared with superforecasters. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads ForecastBench?
o3 leads ForecastBench with 62.5% as of 19 Sept 2026, ahead of Claude Opus 4.1 at 62.0%.
How many models have ForecastBench results?
73 model configurations have a ForecastBench result on BenchLeader, all taken from ForecastBench via Epoch AI Benchmarking Hub.
Who runs ForecastBench and how often is it updated?
ForecastBench is published by ForecastBench. BenchLeader re-reads the published results every morning and records the date each result was published.
Does ForecastBench count toward the BenchLeader Index?
No. ForecastBench is shown for reference but left out of the composite index.