BenchLeader

WinoGrande

Pronoun-resolution commonsense questions. A 2019 benchmark, saturated for current models; kept for history.

As of 19 Sept 2026, Llama 3.1 405B leads WinoGrande on BenchLeader with 89.2%, ahead of Claude 3 Opus at 88.5%, across 74 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Reasoning
Index weight
Reference only
Models
74
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Pronoun-resolution commonsense questions.

How it is scored

Accuracy, as compiled by Epoch AI from published results.

What to keep in mind

Saturated: frontier models score near the ceiling, so it separates only older or smaller models, and results are largely self-reported by developers.

74 of 74
#
1Llama 3.1 405BMetaopen89.2%43.52024-07-23
2Claude 3 OpusAnthropic88.5%39.62024-02-29
3GPT 4 0314OpenAI87.5%43.92023-03-14
4Falcon-180BTII87.1%2023-09-06
5Deepseek v2DeepSeekopen ↗86.3%2024-05-07
6DeepSeek V3DeepSeekopen85.2%44.82024-12-26
7Palm 540BGoogle85.1%2022-04-04
8DeepSeek-Coder-V2 236BDeepSeek83.7%2024-06-17
9Llama 3-70BMetaopen83.5%39.22024-04-18
10PaLM 2-LGoogle83.0%2023-05-17
11Qwen2.5 72BAlibabaopen82.3%42.72024-09-19
12davinci-002OpenAI81.6%2022-03-15
13GPT-3.5 Turbo (older v0613)OpenAI81.6%2023-06-13
14Phi 3 SmallMicrosoft81.5%2024-04-23
15phi-3-medium 14BmediumMicrosoft81.5%2024-04-23
16Qwen2.5-Coder 32B InstructAlibabaopen ↗80.8%42.62024-09-18
17Llama 2-70BMetaopen ↗80.2%34.92023-07-18
18PaLM 2-MGoogle79.2%2023-05-17
19Gemma 7BGoogle79.0%2024-02-21
20Megatron-Turing NLG 530BNVIDIA78.9%2021-10-11
21Falcon 2 11BTII78.3%2024-05-09
22Nemotron-4 15BNVIDIA78.0%2024-02-26
23PaLM 2-SGoogle77.9%2023-05-17
24InstructGPT 175BOpenAI77.7%2022-01-27
25Mixtral 8x7BMistral AIopen77.2%32.12023-12-11
26LLaMA-65BMetaopen77.0%2023-02-24
27PaLM 62BGoogle77.0%2022-04-04
28Falcon-40BTII76.9%2023-03-15
29Qwen2.5-Coder-14BAlibaba76.8%2024-09-18
30Llama 2-34BMeta76.7%2023-07-18
31LLaMA-33BMeta76.0%2023-02-24
32Llama 3-8BMetaopen75.7%35.62024-04-18
33Mistral 7B v0.3Mistral AIopen75.3%32.12023-09-27
34Claude 3 SonnetAnthropic75.1%38.62024-02-29
35Claude 3 HaikuAnthropic74.2%37.62024-03-07
36Phi-1.5Microsoft73.4%2023-09-11
37LLaMA-13BMeta73.0%2023-02-24
38Yi-9B01.AI73.0%2024-03-01
39Qwen2.5-Coder 7B InstructAlibabaopen ↗72.9%2024-09-18
40DeepSeek-Coder-V2-Lite-BaseDeepSeek72.9%2024-06-13
41Llama 2-13BMetaopen ↗72.8%36.52023-07-18
42Yi 6B01.AI71.3%2023-11-02
43MPT-30BMosaicML71.0%2023-06-22
44Phi 3 MiniMicrosoftopen70.8%36.22024-04-23
45vicuna-13b-v1.1LMSYS70.8%2023-04-12
46LLaMA-7BMeta70.1%2023-02-24
47Llama 2-7BMetaopen ↗69.2%34.22023-07-18
48GPT 3.5 Turbo 1106OpenAI68.8%39.22023-11-06
49MPT-7BMosaicML68.6%2023-05-05
50Qwen2.5-Coder-3BAlibaba67.4%2024-09-18
51Falcon-7BTII67.2%2023-04-24
52open_llama_7bUnknown67.0%2023-06-07
53GPT-NeoX-20BEleutherAI66.1%2022-04-07
54INTELLECT-1Prime Intellect65.8%2024-11-29
55Gemma 2BGoogle65.4%2024-02-21
56XGen-7BSalesforce64.9%2023-06-27
57opt-13bUnknown64.7%2022-05-11
58GPT-J-6BEleutherAI64.5%2021-08-05
59StarCoder 2 15BNVIDIA64.3%2024-02-20
60RedPajama-INCITE-7B-BaseUnknown63.8%2023-05-04

Cite as: BenchLeader, “WinoGrande leaderboard”, https://www.benchleader.com/benchmarks/winogrande, data as of 19 Sept 2026.

WinoGrande: questions

What does WinoGrande measure?
Pronoun-resolution commonsense questions. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads WinoGrande?
Llama 3.1 405B leads WinoGrande with 89.2% as of 19 Sept 2026, ahead of Claude 3 Opus at 88.5%.
How many models have WinoGrande results?
74 model configurations have a WinoGrande result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs WinoGrande and how often is it updated?
WinoGrande is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does WinoGrande count toward the BenchLeader Index?
No. WinoGrande is shown for reference but left out of the composite index.