BenchLeader

SciPredict

Predicting the outcomes of scientific experiments before they are run. Scale AI.

As of 19 Sept 2026, Gemini 3 Pro leads SciPredict on BenchLeader with 25.3%, ahead of Claude Opus 4.5 at 23.1%, across 15 model configurations with a published result.

Published by
Scale AI SEAL
Category
Reasoning
Index weight
Reference only
Models
15
Data as of
19 Sept 2026

Scale AI SEAL Leaderboards.

What the test looks like

Given an experimental setup from the literature, the model predicts the result; scored against what the paper found.

How it is scored

Accuracy, published by Scale AI.

What to keep in mind

Tests scientific intuition; contamination is possible for published experiments.

15 of 15
#
1Gemini 3 ProGoogle25.3%61.12026-01-14
2Claude Opus 4.5Anthropic23.1%58.22026-01-14
3Claude Sonnet 4.5Anthropic22.6%54.32026-01-14
4Gemini 3 FlashGoogle22.2%57.72026-01-14
5Claude Opus 4.1Anthropic22.2%51.92026-01-14
6GPT-5.2OpenAI20.6%58.42026-01-14
7o3-miniOpenAI19.8%50.42026-01-14
8DeepSeek V3DeepSeekopen19.2%44.82026-01-14
9Llama 3.3 70BMetaopen18.2%41.42026-01-14
10o3OpenAI17.9%60.32026-01-14
11Gemini 2.5 ProGoogle17.0%54.52026-01-14
12Qwen3 32BAlibabaopen17.0%52.12026-01-14
13Qwen3 235B A22BAlibabaopen16.6%49.62026-01-14
14o4-miniOpenAI16.2%57.42026-01-14
15Llama 3.1 8BMetaopen14.7%36.82026-01-14

Cite as: BenchLeader, “SciPredict leaderboard”, https://www.benchleader.com/benchmarks/scale_scipredict, data as of 19 Sept 2026.

SciPredict: questions

What does SciPredict measure?
Given an experimental setup from the literature, the model predicts the result; scored against what the paper found. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads SciPredict?
Gemini 3 Pro leads SciPredict with 25.3% as of 19 Sept 2026, ahead of Claude Opus 4.5 at 23.1%.
How many models have SciPredict results?
15 model configurations have a SciPredict result on BenchLeader, all taken from Scale AI SEAL.
Who runs SciPredict and how often is it updated?
SciPredict is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
Does SciPredict count toward the BenchLeader Index?
No. SciPredict is shown for reference but left out of the composite index.