WinoGrande
Pronoun-resolution commonsense questions. A 2019 benchmark, saturated for current models; kept for history.
As of 19 Sept 2026, Llama 3.1 405B leads WinoGrande on BenchLeader with 89.2%, ahead of Claude 3 Opus at 88.5%, across 74 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Reasoning
- Index weight
- Reference only
- Models
- 74
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Pronoun-resolution commonsense questions.
How it is scored
Accuracy, as compiled by Epoch AI from published results.
What to keep in mind
Saturated: frontier models score near the ceiling, so it separates only older or smaller models, and results are largely self-reported by developers.
- 1Llama 3.1 405B89.2%
- 2Claude 3 Opus88.5%
- 3GPT 4 031487.5%
- 4Falcon-180B87.1%
- 5Deepseek v286.3%
- 6DeepSeek V385.2%
- 7Palm 540B85.1%
- 8DeepSeek-Coder-V2 236B83.7%
- 9Llama 3-70B83.5%
- 10PaLM 2-L83.0%
- 11Qwen2.5 72B82.3%
- 12davinci-00281.6%
- 13GPT-3.5 Turbo (older v0613)81.6%
- 14Phi 3 Small81.5%
- 15phi-3-medium 14B (medium)81.5%
| # | ||||
|---|---|---|---|---|
| 1 | 89.2% | 43.5 | 2024-07-23 | |
| 2 | 88.5% | 39.6 | 2024-02-29 | |
| 3 | 87.5% | 43.9 | 2023-03-14 | |
| 4 | 87.1% | – | 2023-09-06 | |
| 5 | 86.3% | – | 2024-05-07 | |
| 6 | 85.2% | 44.8 | 2024-12-26 | |
| 7 | 85.1% | – | 2022-04-04 | |
| 8 | 83.7% | – | 2024-06-17 | |
| 9 | 83.5% | 39.2 | 2024-04-18 | |
| 10 | 83.0% | – | 2023-05-17 | |
| 11 | 82.3% | 42.7 | 2024-09-19 | |
| 12 | 81.6% | – | 2022-03-15 | |
| 13 | 81.6% | – | 2023-06-13 | |
| 14 | 81.5% | – | 2024-04-23 | |
| 15 | 81.5% | – | 2024-04-23 | |
| 16 | 80.8% | 42.6 | 2024-09-18 | |
| 17 | 80.2% | 34.9 | 2023-07-18 | |
| 18 | 79.2% | – | 2023-05-17 | |
| 19 | 79.0% | – | 2024-02-21 | |
| 20 | 78.9% | – | 2021-10-11 | |
| 21 | 78.3% | – | 2024-05-09 | |
| 22 | 78.0% | – | 2024-02-26 | |
| 23 | 77.9% | – | 2023-05-17 | |
| 24 | 77.7% | – | 2022-01-27 | |
| 25 | 77.2% | 32.1 | 2023-12-11 | |
| 26 | 77.0% | – | 2023-02-24 | |
| 27 | 77.0% | – | 2022-04-04 | |
| 28 | 76.9% | – | 2023-03-15 | |
| 29 | 76.8% | – | 2024-09-18 | |
| 30 | 76.7% | – | 2023-07-18 | |
| 31 | 76.0% | – | 2023-02-24 | |
| 32 | 75.7% | 35.6 | 2024-04-18 | |
| 33 | 75.3% | 32.1 | 2023-09-27 | |
| 34 | 75.1% | 38.6 | 2024-02-29 | |
| 35 | 74.2% | 37.6 | 2024-03-07 | |
| 36 | 73.4% | – | 2023-09-11 | |
| 37 | 73.0% | – | 2023-02-24 | |
| 38 | 73.0% | – | 2024-03-01 | |
| 39 | 72.9% | – | 2024-09-18 | |
| 40 | 72.9% | – | 2024-06-13 | |
| 41 | 72.8% | 36.5 | 2023-07-18 | |
| 42 | 71.3% | – | 2023-11-02 | |
| 43 | 71.0% | – | 2023-06-22 | |
| 44 | 70.8% | 36.2 | 2024-04-23 | |
| 45 | 70.8% | – | 2023-04-12 | |
| 46 | 70.1% | – | 2023-02-24 | |
| 47 | 69.2% | 34.2 | 2023-07-18 | |
| 48 | 68.8% | 39.2 | 2023-11-06 | |
| 49 | 68.6% | – | 2023-05-05 | |
| 50 | 67.4% | – | 2024-09-18 | |
| 51 | 67.2% | – | 2023-04-24 | |
| 52 | open_llama_7bUnknown | 67.0% | – | 2023-06-07 |
| 53 | GPT-NeoX-20BEleutherAI | 66.1% | – | 2022-04-07 |
| 54 | 65.8% | – | 2024-11-29 | |
| 55 | 65.4% | – | 2024-02-21 | |
| 56 | 64.9% | – | 2023-06-27 | |
| 57 | opt-13bUnknown | 64.7% | – | 2022-05-11 |
| 58 | GPT-J-6BEleutherAI | 64.5% | – | 2021-08-05 |
| 59 | 64.3% | – | 2024-02-20 | |
| 60 | RedPajama-INCITE-7B-BaseUnknown | 63.8% | – | 2023-05-04 |
Cite as: BenchLeader, “WinoGrande leaderboard”, https://www.benchleader.com/benchmarks/winogrande, data as of 19 Sept 2026.
WinoGrande: questions
- What does WinoGrande measure?
- Pronoun-resolution commonsense questions. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads WinoGrande?
- Llama 3.1 405B leads WinoGrande with 89.2% as of 19 Sept 2026, ahead of Claude 3 Opus at 88.5%.
- How many models have WinoGrande results?
- 74 model configurations have a WinoGrande result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs WinoGrande and how often is it updated?
- WinoGrande is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does WinoGrande count toward the BenchLeader Index?
- No. WinoGrande is shown for reference but left out of the composite index.