VPCT
Visual physics comprehension: predicting what happens next in a scene.
As of 19 Sept 2026, Gemini 3 Pro leads VPCT on BenchLeader with 91.0%, ahead of GPT-5.2 at 84.0%, across 30 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Multimodal
- Index weight
- Reference only
- Models
- 30
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Images of physical setups; the model predicts the outcome.
How it is scored
Percent correct, as published.
What to keep in mind
Vision models only; small set.
- 1Gemini 3 Pro91.0%
- 2GPT-5.2 (xhigh)84.0%
- 3Gemini 3 Flash72.6%
- 4GPT-5.2 (high)67.0%
- 5GPT-5 (high)66.0%
- 6GPT-5 (medium)63.2%
- 7GPT-5.1 (high)58.7%
- 8o4-mini (medium)57.5%
- 9GPT-5.1 (medium)53.3%
- 10o3 (medium)52.0%
- 11Gemini 2.5 Pro48.0%
- 12Gemini 2.5 Flash 09 202546.2%
- 13GPT-4.545.0%
- 14gemini-robotics-er-1.5-preview40.8%
- 15GPT-5 mini (medium)40.2%
30 of 30
| # | ||||
|---|---|---|---|---|
| 1 | 91.0% | 61.1 | 2025-11-18 | |
| 2 | 84.0% | 62.0 | 2025-12-11 | |
| 3 | 72.6% | 57.7 | 2025-12-17 | |
| 4 | 67.0% | 55.7 | 2025-12-11 | |
| 5 | 66.0% | 58.6 | 2025-08-07 | |
| 6 | 63.2% | 58.5 | 2025-08-07 | |
| 7 | 58.7% | 58.7 | 2025-11-13 | |
| 8 | 57.5% | 52.4 | 2025-04-16 | |
| 9 | 53.3% | 52.8 | 2025-11-13 | |
| 10 | 52.0% | 53.7 | 2025-04-16 | |
| 11 | 48.0% | 54.5 | 2025-03-31 | |
| 12 | 46.2% | 52.1 | 2025-09-25 | |
| 13 | 45.0% | 50.2 | 2025-02-27 | |
| 14 | 40.8% | – | 2025-09-26 | |
| 15 | 40.2% | 54.6 | 2025-08-07 | |
| 16 | 40.0% | 58.2 | 2025-11-24 | |
| 17 | 40.0% | 44.3 | 2024-11-20 | |
| 18 | 39.8% | 54.3 | 2025-09-29 | |
| 19 | 39.0% | 55.3 | 2025-08-07 | |
| 20 | 39.0% | 50.2 | 2025-02-24 | |
| 21 | 38.0% | 52.9 | 2025-05-22 | |
| 22 | 38.0% | 52.5 | 2025-04-17 | |
| 23 | 37.2% | 48.2 | 2025-08-07 | |
| 24 | 37.0% | 48.4 | 2024-12-17 | |
| 25 | 35.4% | 46.5 | 2025-08-07 | |
| 26 | 35.0% | 51.9 | 2025-08-05 | |
| 27 | 34.0% | 49.7 | 2025-05-22 | |
| 28 | 34.0% | 36.2 | 2024-07-18 | |
| 29 | 33.0% | 46.6 | 2024-10-22 | |
| 30 | 30.0% | 46.9 | 2025-09-25 |
Cite as: BenchLeader, “VPCT leaderboard”, https://www.benchleader.com/benchmarks/vpct, data as of 19 Sept 2026.
VPCT: questions
- What does VPCT measure?
- Images of physical setups; the model predicts the outcome. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads VPCT?
- Gemini 3 Pro leads VPCT with 91.0% as of 19 Sept 2026, ahead of GPT-5.2 at 84.0%.
- How many models have VPCT results?
- 30 model configurations have a VPCT result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs VPCT and how often is it updated?
- VPCT is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does VPCT count toward the BenchLeader Index?
- No. VPCT is shown for reference but left out of the composite index.