CritPt
Unpublished physics research problems at the level of a first PhD project. Run by Artificial Analysis.
As of 19 Sept 2026, GPT-5.6 Sol leads CritPt on BenchLeader with 32.3%, ahead of GPT-6 Astra at 31.7%, across 496 model configurations with a published result.
- Published by
- Artificial Analysis
- Category
- Reasoning
- Index weight
- 1.0
- Models
- 496
- Data as of
- 19 Sept 2026
Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.
What the test looks like
CritPt poses original research-level physics problems, written by physicists to take a strong graduate student days, with answers checked against a reference. The problems are unpublished so they cannot be memorised.
How it is scored
Accuracy on the problems, run by Artificial Analysis in a fixed harness.
What to keep in mind
Scores are low for every model and small differences are noise; it is a frontier signal rather than a ranking of everyday physics knowledge. Counts toward the reasoning category since index v1.2.
- 1GPT-5.6 Sol (max)32.3%
- 2GPT-6 Astra (max)31.7%
- 3GPT-6 Astra (xhigh)31.4%
- 4Claude Fable 5.1 (xhigh)31.1%
- 5GPT-5.5 Pro (xhigh)30.6%
- 6Claude Fable 5.1 (high)30.3%
- 7GPT-5.6 Terra (max)30.0%
- 8GPT-5.4 Pro (xhigh)30.0%
- 9Claude Fable 5.1 (thinking)29.7%
- 10Claude Opus 5 (max)29.1%
- 11GPT-6 Astra (medium)29.1%
- 12Claude Fable 5.1 (medium)29.1%
- 13GPT-6 Astra (high)28.9%
- 14Claude Fable 5 (thinking)28.6%
- 15GPT-5.6 Sol (xhigh)28.6%
| # | |||
|---|---|---|---|
| 1 | 32.3% | 68.8 | |
| 2 | 31.7% | 71.8 | |
| 3 | 31.4% | 71.1 | |
| 4 | 31.1% | 71.6 | |
| 5 | 30.6% | 67.1 | |
| 6 | 30.3% | 72.0 | |
| 7 | 30.0% | 65.0 | |
| 8 | 30.0% | 63.3 | |
| 9 | 29.7% | 71.3 | |
| 10 | 29.1% | 69.9 | |
| 11 | 29.1% | 69.8 | |
| 12 | 29.1% | 69.6 | |
| 13 | 28.9% | 71.9 | |
| 14 | 28.6% | 70.9 | |
| 15 | 28.6% | 68.2 | |
| 16 | 28.3% | 70.2 | |
| 17 | 27.7% | 70.2 | |
| 18 | 27.7% | 67.5 | |
| 19 | 27.1% | 67.6 | |
| 20 | 27.1% | 64.3 | |
| 21 | 26.9% | 67.3 | |
| 22 | 26.3% | 68.2 | |
| 23 | 26.0% | 68.0 | |
| 24 | 25.7% | 68.0 | |
| 25 | 25.7% | – | |
| 26 | 25.4% | 67.0 | |
| 27 | 24.9% | 69.3 | |
| 28 | 23.4% | 67.2 | |
| 29 | 23.4% | 65.4 | |
| 30 | 23.1% | 63.4 | |
| 31 | 22.9% | 66.1 | |
| 32 | 22.9% | 62.8 | |
| 33 | New | 20.9% | 65.0 |
| 34 | 20.9% | 64.2 | |
| 35 | 20.9% | 63.9 | |
| 36 | 20.6% | 61.1 | |
| 37 | 20.6% | 60.2 | |
| 38 | 20.0% | 66.5 | |
| 39 | 20.0% | 65.2 | |
| 40 | 19.7% | 64.6 | |
| 41 | 19.1% | 65.8 | |
| 42 | 18.6% | 64.6 | |
| 43 | 18.3% | 64.5 | |
| 44 | 18.0% | 64.2 | |
| 45 | 17.7% | 66.2 | |
| 46 | 17.7% | 64.1 | |
| 47 | 17.7% | 63.9 | |
| 48 | 17.4% | 58.4 | |
| 49 | 17.1% | 64.5 | |
| 50 | 16.9% | 65.4 | |
| 51 | 16.9% | 60.5 | |
| 52 | 16.6% | 62.1 | |
| 53 | 16.6% | 58.6 | |
| 54 | 15.7% | 60.8 | |
| 55 | 15.4% | 63.9 | |
| 56 | 15.4% | 61.0 | |
| 57 | 15.4% | 59.8 | |
| 58 | 15.1% | 62.3 | |
| 59 | 15.1% | 62.2 | |
| 60 | 15.1% | 61.8 |
Cite as: BenchLeader, “CritPt leaderboard”, https://www.benchleader.com/benchmarks/aa_critpt, data as of 19 Sept 2026.
CritPt: questions
- What does CritPt measure?
- CritPt poses original research-level physics problems, written by physicists to take a strong graduate student days, with answers checked against a reference. The problems are unpublished so they cannot be memorised. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads CritPt?
- GPT-5.6 Sol leads CritPt with 32.3% as of 19 Sept 2026, ahead of GPT-6 Astra at 31.7%.
- How many models have CritPt results?
- 496 model configurations have a CritPt result on BenchLeader, all taken from Artificial Analysis.
- Who runs CritPt and how often is it updated?
- CritPt is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does CritPt count toward the BenchLeader Index?
- Yes. CritPt contributes to the reasoning category of the BenchLeader Index, normalised so that 50 is the average of the evaluated models.