BenchLeader

VPCT

Visual physics comprehension: predicting what happens next in a scene.

As of 19 Sept 2026, Gemini 3 Pro leads VPCT on BenchLeader with 91.0%, ahead of GPT-5.2 at 84.0%, across 30 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Multimodal
Index weight
Reference only
Models
30
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Images of physical setups; the model predicts the outcome.

How it is scored

Percent correct, as published.

What to keep in mind

Vision models only; small set.

30 of 30
#
1Gemini 3 ProGoogle91.0%61.12025-11-18
2GPT-5.2xhighOpenAI84.0%62.02025-12-11
3Gemini 3 FlashGoogle72.6%57.72025-12-17
4GPT-5.2highOpenAI67.0%55.72025-12-11
5GPT-5highOpenAI66.0%58.62025-08-07
6GPT-5mediumOpenAI63.2%58.52025-08-07
7GPT-5.1highOpenAI58.7%58.72025-11-13
8o4-minimediumOpenAI57.5%52.42025-04-16
9GPT-5.1mediumOpenAI53.3%52.82025-11-13
10o3mediumOpenAI52.0%53.72025-04-16
11Gemini 2.5 ProGoogle48.0%54.52025-03-31
12Gemini 2.5 Flash 09 2025Google46.2%52.12025-09-25
13GPT-4.5OpenAI45.0%50.22025-02-27
14gemini-robotics-er-1.5-previewGoogle40.8%2025-09-26
15GPT-5 minimediumOpenAI40.2%54.62025-08-07
16Claude Opus 4.5Anthropic40.0%58.22025-11-24
17GPT-4oOpenAI40.0%44.32024-11-20
18Claude Sonnet 4.5Anthropic39.8%54.32025-09-29
19GPT-5 minihighOpenAI39.0%55.32025-08-07
20Claude 3.7 SonnetAnthropic39.0%50.22025-02-24
21Claude Opus 4Anthropic38.0%52.92025-05-22
22Gemini 2.5 FlashGoogle38.0%52.52025-04-17
23GPT-5 nanohighOpenAI37.2%48.22025-08-07
24o1mediumOpenAI37.0%48.42024-12-17
25GPT-5 nanomediumOpenAI35.4%46.52025-08-07
26Claude Opus 4.1Anthropic35.0%51.92025-08-05
27Claude Sonnet 4Anthropic34.0%49.72025-05-22
28GPT-4o miniOpenAI34.0%36.22024-07-18
29Claude 3.5 SonnetAnthropic33.0%46.62024-10-22
30Gemini 2.5 Flash Lite 09Google30.0%46.92025-09-25

Cite as: BenchLeader, “VPCT leaderboard”, https://www.benchleader.com/benchmarks/vpct, data as of 19 Sept 2026.

VPCT: questions

What does VPCT measure?
Images of physical setups; the model predicts the outcome. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads VPCT?
Gemini 3 Pro leads VPCT with 91.0% as of 19 Sept 2026, ahead of GPT-5.2 at 84.0%.
How many models have VPCT results?
30 model configurations have a VPCT result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs VPCT and how often is it updated?
VPCT is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does VPCT count toward the BenchLeader Index?
No. VPCT is shown for reference but left out of the composite index.