SciPredict
Predicting the outcomes of scientific experiments before they are run. Scale AI.
As of 19 Sept 2026, Gemini 3 Pro leads SciPredict on BenchLeader with 25.3%, ahead of Claude Opus 4.5 at 23.1%, across 15 model configurations with a published result.
- Published by
- Scale AI SEAL
- Category
- Reasoning
- Index weight
- Reference only
- Models
- 15
- Data as of
- 19 Sept 2026
Scale AI SEAL Leaderboards.
What the test looks like
Given an experimental setup from the literature, the model predicts the result; scored against what the paper found.
How it is scored
Accuracy, published by Scale AI.
What to keep in mind
Tests scientific intuition; contamination is possible for published experiments.
- 1Gemini 3 Pro25.3%
- 2Claude Opus 4.523.1%
- 3Claude Sonnet 4.522.6%
- 4Gemini 3 Flash22.2%
- 5Claude Opus 4.122.2%
- 6GPT-5.220.6%
- 7o3-mini19.8%
- 8DeepSeek V319.2%
- 9Llama 3.3 70B18.2%
- 10o317.9%
- 11Gemini 2.5 Pro17.0%
- 12Qwen3 32B17.0%
- 13Qwen3 235B A22B16.6%
- 14o4-mini16.2%
- 15Llama 3.1 8B14.7%
15 of 15
| # | ||||
|---|---|---|---|---|
| 1 | 25.3% | 61.1 | 2026-01-14 | |
| 2 | 23.1% | 58.2 | 2026-01-14 | |
| 3 | 22.6% | 54.3 | 2026-01-14 | |
| 4 | 22.2% | 57.7 | 2026-01-14 | |
| 5 | 22.2% | 51.9 | 2026-01-14 | |
| 6 | 20.6% | 58.4 | 2026-01-14 | |
| 7 | 19.8% | 50.4 | 2026-01-14 | |
| 8 | 19.2% | 44.8 | 2026-01-14 | |
| 9 | 18.2% | 41.4 | 2026-01-14 | |
| 10 | 17.9% | 60.3 | 2026-01-14 | |
| 11 | 17.0% | 54.5 | 2026-01-14 | |
| 12 | 17.0% | 52.1 | 2026-01-14 | |
| 13 | 16.6% | 49.6 | 2026-01-14 | |
| 14 | 16.2% | 57.4 | 2026-01-14 | |
| 15 | 14.7% | 36.8 | 2026-01-14 |
Cite as: BenchLeader, “SciPredict leaderboard”, https://www.benchleader.com/benchmarks/scale_scipredict, data as of 19 Sept 2026.
SciPredict: questions
- What does SciPredict measure?
- Given an experimental setup from the literature, the model predicts the result; scored against what the paper found. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads SciPredict?
- Gemini 3 Pro leads SciPredict with 25.3% as of 19 Sept 2026, ahead of Claude Opus 4.5 at 23.1%.
- How many models have SciPredict results?
- 15 model configurations have a SciPredict result on BenchLeader, all taken from Scale AI SEAL.
- Who runs SciPredict and how often is it updated?
- SciPredict is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does SciPredict count toward the BenchLeader Index?
- No. SciPredict is shown for reference but left out of the composite index.