PIQA
Physical commonsense questions. A 2019 benchmark, saturated for current models; kept for history.
As of 19 Sept 2026, GPT-4o mini leads PIQA on BenchLeader with 88.7%, ahead of Phi-3.5-MoE at 88.6%, across 56 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Reasoning
- Index weight
- Reference only
- Models
- 56
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Physical commonsense questions.
How it is scored
Accuracy, as compiled by Epoch AI from published results.
What to keep in mind
Saturated: frontier models score near the ceiling, so it separates only older or smaller models, and results are largely self-reported by developers.
- 1GPT-4o mini88.7%
- 2Phi-3.5-MoE88.6%
- 3Gemini 1.5 Flash 00287.5%
- 4Llama 3.1 405B85.9%
- 5Falcon-180B84.9%
- 6DeepSeek-V384.7%
- 7Inflection-184.2%
- 8Deepseek v283.9%
- 9Gemma 2 9B83.7%
- 10Mixtral 8x7B83.6%
- 11Mistral NeMo83.5%
- 12Stable Beluga 283.3%
- 13Megatron-Turing NLG 530B83.2%
- 14Mistral 7B v0.383.0%
- 15Falcon-40B83.0%
56 of 56
| # | ||||
|---|---|---|---|---|
| 1 | 88.7% | 36.2 | 2024-07-18 | |
| 2 | 88.6% | – | 2024-08-17 | |
| 3 | 87.5% | 40.8 | 2024-09-24 | |
| 4 | 85.9% | 43.5 | 2024-07-23 | |
| 5 | 84.9% | – | 2023-09-06 | |
| 6 | 84.7% | – | 2024-12-26 | |
| 7 | Inflection-1Inflection AI | 84.2% | – | 2023-06-22 |
| 8 | 83.9% | – | 2024-05-07 | |
| 9 | 83.7% | 39.5 | – | |
| 10 | 83.6% | 32.1 | 2023-12-11 | |
| 11 | 83.5% | – | 2024-07-18 | |
| 12 | 83.3% | – | 2023-07-20 | |
| 13 | 83.2% | – | 2021-10-11 | |
| 14 | 83.0% | 32.1 | 2023-09-27 | |
| 15 | 83.0% | – | 2023-03-15 | |
| 16 | 82.8% | 34.9 | 2023-07-18 | |
| 17 | 82.8% | – | 2023-02-24 | |
| 18 | 82.6% | 42.7 | 2024-09-19 | |
| 19 | 82.4% | – | 2024-02-26 | |
| 20 | 82.3% | – | 2022-01-27 | |
| 21 | 82.3% | – | 2023-02-24 | |
| 22 | 82.3% | – | 2022-04-04 | |
| 23 | 81.9% | – | 2023-06-22 | |
| 24 | 81.9% | – | 2023-07-18 | |
| 25 | 81.2% | 36.8 | 2024-07-23 | |
| 26 | 81.2% | – | 2024-02-21 | |
| 27 | 81.0% | – | 2024-08-16 | |
| 28 | 80.8% | 36.5 | 2023-07-18 | |
| 29 | 80.6% | – | 2023-05-05 | |
| 30 | 80.5% | – | 2022-04-04 | |
| 31 | 80.3% | – | 2023-04-24 | |
| 32 | 80.3% | – | 2023-09-18 | |
| 33 | 80.1% | – | 2023-02-24 | |
| 34 | 79.9% | – | 2023-09-28 | |
| 35 | 79.8% | – | 2023-02-24 | |
| 36 | 78.8% | 34.2 | 2023-07-18 | |
| 37 | 78.1% | – | 2023-09-06 | |
| 38 | 77.9% | – | 2023-07-05 | |
| 39 | 77.9% | – | 2023-09-28 | |
| 40 | 77.4% | – | 2023-04-12 | |
| 41 | 77.3% | – | 2024-02-21 | |
| 42 | RedPajama-INCITE-7B-BaseUnknown | 76.9% | – | 2023-05-04 |
| 43 | GPT-NeoX-20BEleutherAI | 76.7% | – | 2022-04-07 |
| 44 | 76.2% | – | 2023-06-01 | |
| 45 | open_llama_7bUnknown | 76.0% | – | 2023-06-07 |
| 46 | opt-13bUnknown | 75.7% | – | 2022-05-11 |
| 47 | 75.5% | – | 2023-06-27 | |
| 48 | 75.4% | – | 2023-04-11 | |
| 49 | GPT-J-6BEleutherAI | 75.4% | – | 2021-08-05 |
| 50 | 73.5% | – | 2023-03-20 | |
| 51 | 73.3% | – | 2023-11-30 | |
| 52 | GPT-Neo-2.7BEleutherAI | 72.9% | – | 2021-03-30 |
| 53 | 70.5% | – | 2019-11-05 | |
| 54 | 69.6% | – | 2023-06-24 | |
| 55 | 69.0% | – | 2022-05-11 | |
| 56 | 65.8% | – | 2023-04-19 |
Cite as: BenchLeader, “PIQA leaderboard”, https://www.benchleader.com/benchmarks/piqa, data as of 19 Sept 2026.
PIQA: questions
- What does PIQA measure?
- Physical commonsense questions. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads PIQA?
- GPT-4o mini leads PIQA with 88.7% as of 19 Sept 2026, ahead of Phi-3.5-MoE at 88.6%.
- How many models have PIQA results?
- 56 model configurations have a PIQA result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs PIQA and how often is it updated?
- PIQA is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does PIQA count toward the BenchLeader Index?
- No. PIQA is shown for reference but left out of the composite index.