BenchLeader

CritPt

Unpublished physics research problems at the level of a first PhD project. Run by Artificial Analysis.

As of 19 Sept 2026, GPT-5.6 Sol leads CritPt on BenchLeader with 32.3%, ahead of GPT-6 Astra at 31.7%, across 496 model configurations with a published result.

Published by
Artificial Analysis
Category
Reasoning
Index weight
1.0
Models
496
Data as of
19 Sept 2026

Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.

What the test looks like

CritPt poses original research-level physics problems, written by physicists to take a strong graduate student days, with answers checked against a reference. The problems are unpublished so they cannot be memorised.

How it is scored

Accuracy on the problems, run by Artificial Analysis in a fixed harness.

What to keep in mind

Scores are low for every model and small differences are noise; it is a frontier signal rather than a ranking of everyday physics knowledge. Counts toward the reasoning category since index v1.2.

496 of 496
#
1GPT-5.6 SolmaxOpenAI32.3%68.8
2GPT-6 AstramaxOpenAI31.7%71.8
3GPT-6 AstraxhighOpenAI31.4%71.1
4Claude Fable 5.1xhighAnthropic31.1%71.6
5GPT-5.5 ProxhighOpenAI30.6%67.1
6Claude Fable 5.1highAnthropic30.3%72.0
7GPT-5.6 TerramaxOpenAI30.0%65.0
8GPT-5.4 ProxhighOpenAI30.0%63.3
9Claude Fable 5.1thinkingAnthropic29.7%71.3
10Claude Opus 5maxAnthropic29.1%69.9
11GPT-6 AstramediumOpenAI29.1%69.8
12Claude Fable 5.1mediumAnthropic29.1%69.6
13GPT-6 AstrahighOpenAI28.9%71.9
14Claude Fable 5thinkingAnthropic28.6%70.9
15GPT-5.6 SolxhighOpenAI28.6%68.2
16Claude Opus 5highAnthropic28.3%70.2
17Claude Opus 5xhighAnthropic27.7%70.2
18Claude Fable 5.1lowAnthropic27.7%67.5
19GPT-5.5xhighOpenAI27.1%67.6
20GPT-5.6 TerraxhighOpenAI27.1%64.3
21Claude Opus 5mediumAnthropic26.9%67.3
22GPT-6 AstralowOpenAI26.3%68.2
23Muse Spark 1.3xhighMeta26.0%68.0
24GPT-5.6 SolhighOpenAI25.7%68.0
25Gemini 3 Deep ThinkthinkingGoogle25.7%
26GPT-5.5highOpenAI25.4%67.0
27Muse Spark 1.3maxMeta24.9%69.3
28Kimi K3maxMoonshot AIopen ↗23.4%67.2
29GPT-5.4xhighOpenAI23.4%65.4
30Claude Opus 5lowAnthropic23.1%63.4
31GPT-5.6 SolmediumOpenAI22.9%66.1
32GPT-5.6 TerrahighOpenAI22.9%62.8
33NewStep 5 PreviewStepFun20.9%65.0
34Claude Opus 4.8maxAnthropic20.9%64.2
35GLM-5.2maxZhipu AIopen ↗20.9%63.9
36GPT-5.6 LunaxhighOpenAI20.6%61.1
37GPT-5.6 LunamaxOpenAI20.6%60.2
38Qwen3 8maxAlibaba20.0%66.5
39Qwen3.8 2.4T A95BAlibabaopen ↗20.0%65.2
40Grok 4.6xhighxAI19.7%64.6
41GLM-5.3maxZhipu AIopen ↗19.1%65.8
42GPT-5.5mediumOpenAI18.6%64.6
43Gemini 3.8 FlashhighGoogle18.3%64.5
44DeepSeek V4 PromaxDeepSeekopen ↗18.0%64.2
45Grok 4.6mediumxAI17.7%66.2
46Muse Spark 1.2xhighMeta17.7%64.1
47Gemini 3.1 ProGoogle17.7%63.9
48GPT-5.6 TerramediumOpenAI17.4%58.4
49Grok 4.6highxAI17.1%64.5
50GPT-5.3 CodexxhighOpenAI16.9%65.4
51Claude Sonnet 5maxAnthropic16.9%60.5
52DeepSeek V4 FlashmaxDeepSeekopen ↗16.6%62.1
53GPT-5.6 LunahighOpenAI16.6%58.6
54Agnes 2.5 Pro BetaSapiens AI15.7%60.8
55GLM-5.3-FlashZhipu AIopen ↗15.4%63.9
56Grok 4.5highxAI15.4%61.0
57Claude Sonnet 5xhighAnthropic15.4%59.8
58Agnes 3.0 FlashSapiens AI15.1%62.3
59Claude Sonnet 5highAnthropic15.1%62.2
60Muse Spark 1.1xhighMeta15.1%61.8

Cite as: BenchLeader, “CritPt leaderboard”, https://www.benchleader.com/benchmarks/aa_critpt, data as of 19 Sept 2026.

CritPt: questions

What does CritPt measure?
CritPt poses original research-level physics problems, written by physicists to take a strong graduate student days, with answers checked against a reference. The problems are unpublished so they cannot be memorised. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads CritPt?
GPT-5.6 Sol leads CritPt with 32.3% as of 19 Sept 2026, ahead of GPT-6 Astra at 31.7%.
How many models have CritPt results?
496 model configurations have a CritPt result on BenchLeader, all taken from Artificial Analysis.
Who runs CritPt and how often is it updated?
CritPt is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
Does CritPt count toward the BenchLeader Index?
Yes. CritPt contributes to the reasoning category of the BenchLeader Index, normalised so that 50 is the average of the evaluated models.