BenchLeader

HellaSwag

Commonsense sentence completion. A 2019 benchmark, saturated for current models; kept for history.

As of 19 Sept 2026, GPT 4 0314 leads HellaSwag on BenchLeader with 95.3%, ahead of Llama 3.1 405B at 89.2%, across 68 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Reasoning
Index weight
Reference only
Models
68
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Commonsense sentence completion.

How it is scored

Accuracy, as compiled by Epoch AI from published results.

What to keep in mind

Saturated: frontier models score near the ceiling, so it separates only older or smaller models, and results are largely self-reported by developers.

68 of 68
#
1GPT 4 0314OpenAI95.3%43.92023-03-14
2Llama 3.1 405BMetaopen89.2%43.52024-07-23
3Falcon-180BTII89.0%2023-09-06
4DeepSeek V3DeepSeekopen88.9%44.82024-12-26
5Deepseek v2DeepSeekopen ↗87.1%2024-05-07
6PaLM 2-LGoogle86.8%2023-05-17
7Mixtral 8x7BMistral AIopen86.7%32.12023-12-11
8davinci-003OpenAI85.5%2022-11-28
9Llama 2-70BMetaopen ↗85.3%34.92023-07-18
10Falcon-40BTII85.3%2023-03-15
11Qwen2.5 72BAlibabaopen84.8%42.72024-09-19
12LLaMA-65BMetaopen84.2%2023-02-24
13Stable Beluga 2Stability AI84.1%2023-07-20
14PaLM 2-MGoogle84.0%2023-05-17
15Palm 540BGoogle83.8%2022-04-04
16Qwen2.5-Coder 32B InstructAlibabaopen ↗83.0%42.62024-09-18
17Falcon 2 11BTII82.9%2024-05-09
18LLaMA-33BMeta82.8%2023-02-24
19phi-3-medium 14BmediumMicrosoft82.4%2024-04-23
20Megatron-Turing NLG 530BNVIDIA82.4%2021-10-11
21Nemotron-4 15BNVIDIA82.4%2024-02-26
22Gemma 7BGoogle82.2%2024-02-21
23PaLM 2-SGoogle82.0%2023-05-17
24davinci-002OpenAI81.5%2022-03-15
25Mistral 7B v0.3Mistral AIopen81.0%32.12023-09-27
26Llama 2-13BMetaopen ↗80.7%36.52023-07-18
27Qwen2.5-Coder-14BAlibaba80.2%2024-09-18
28InstructGPT 175BOpenAI79.3%2022-01-27
29LLaMA-13BMeta79.2%2023-02-24
30OPT-175BMeta79.1%2022-05-02
31Falcon-7BTII78.1%2023-04-24
32internlm-20bShanghai AI Lab78.1%2023-09-18
33Llama 2-7BMetaopen ↗77.2%34.22023-07-18
34Phi 3 SmallMicrosoft77.0%2024-04-23
35Qwen2.5-Coder 7B InstructAlibabaopen ↗76.8%2024-09-18
36Phi 3 MiniMicrosoftopen76.7%36.22024-04-23
37MPT-7BMosaicML76.4%2023-05-05
38Yi-9B01.AI76.4%2024-03-01
39LLaMA-7BMeta76.2%2023-02-24
40OPT-66BMeta74.5%2022-05-03
41Yi 6B01.AI74.4%2023-11-02
42XGen-7BSalesforce74.2%2023-06-27
43open_llama_7bUnknown71.8%2023-06-07
44INTELLECT-1Prime Intellect71.4%2024-11-29
45Gemma 2BGoogle71.4%2024-02-21
46Qwen2.5-Coder-3BAlibaba70.9%2024-09-18
47Dolly v2 12BDatabricks70.8%2023-04-11
48Baichuan2-13BBaichuan70.8%2023-09-06
49internlm-7bShanghai AI Lab70.6%2023-07-05
50GPT-NeoX-20BEleutherAI70.5%2022-04-07
51RedPajama-INCITE-7B-BaseUnknown70.3%2023-05-04
52opt-13bUnknown69.9%2022-05-11
53Baichuan 2-7BBaichuan68.0%2023-09-20
54InstructGPT 6BOpenAI67.6%
55GPT-J-6BEleutherAI66.2%2021-08-05
56Qwen2 5 Coder 1 5BAlibaba61.8%2024-09-18
57Cerebras-GPT-13BCerebras59.4%2023-03-20
58vicuna-13b-v1.1LMSYS57.8%2023-04-12
59chatglm2-6bZhipu AI57.0%2023-06-24
60InstructGPT 1.3BOpenAI56.1%

Cite as: BenchLeader, “HellaSwag leaderboard”, https://www.benchleader.com/benchmarks/hellaswag, data as of 19 Sept 2026.

HellaSwag: questions

What does HellaSwag measure?
Commonsense sentence completion. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads HellaSwag?
GPT 4 0314 leads HellaSwag with 95.3% as of 19 Sept 2026, ahead of Llama 3.1 405B at 89.2%.
How many models have HellaSwag results?
68 model configurations have a HellaSwag result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs HellaSwag and how often is it updated?
HellaSwag is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does HellaSwag count toward the BenchLeader Index?
No. HellaSwag is shown for reference but left out of the composite index.