BenchLeader

GSM8K

Grade-school maths word problems. A 2021 benchmark, saturated for current models; kept for history.

As of 19 Sept 2026, Deepseek Coder v2 leads GSM8K on BenchLeader with 94.5%, ahead of Qwen2.5-Coder-14B at 94.2%, across 74 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Maths
Index weight
Reference only
Models
74
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Grade-school maths word problems.

How it is scored

Accuracy, as compiled by Epoch AI from published results.

What to keep in mind

Saturated: frontier models score near the ceiling, so it separates only older or smaller models, and results are largely self-reported by developers.

74 of 74
#
1Deepseek Coder v2DeepSeekopen ↗94.5%45.42024-06-17
2Qwen2.5-Coder-14BAlibaba94.2%2024-11-06
3Qwen2.5-Coder 32B InstructAlibabaopen ↗93.0%42.62024-11-21
4GPT 4 0314OpenAI92.0%43.92023-03-14
5GPT-4o miniOpenAI91.3%36.22024-07-18
6GPT 4 0613OpenAI90.0%40.22023-06-13
7Phi-3.5-MoEMicrosoft88.7%2024-08-17
8DeepSeek-Coder-V2-Lite-InstructDeepSeekopen87.6%2024-06-13
9Qwen2.5-Coder 7B InstructAlibabaopen ↗86.7%2024-09-17
10Claude InstantAnthropic86.7%2023-08-09
11Phi-3.5-miniMicrosoft86.2%2024-08-16
12DeepSeek-Coder-V2 236BDeepSeek85.8%2024-06-17
13Gemma 2 9BGoogle84.9%39.5
14Mistral NeMoMistral AI84.2%2024-07-18
15Llama 3.1 8BMetaopen82.4%36.82024-07-23
16Gemini 1.5 Flash 001Google82.4%36.72024-05-23
17Claude InstantAnthropic80.9%
18Qwen2.5-Coder-3BAlibaba80.7%2024-11-06
19davinci-003OpenAI78.2%2022-11-28
20Yi-34B01.AI76.0%35.52023-11-22
21Mixtral 8x7BMistral AIopen74.4%32.12023-12-11
22Llama 2-70BMetaopen ↗69.6%34.92023-07-18
23Stable Beluga 2Stability AI69.6%2023-07-20
24DeepSeek-Coder-V2-Lite-BaseDeepSeek67.1%2024-06-13
25Qwen2 5 Coder 1 5BAlibaba65.8%2024-09-18
26internlm-20bShanghai AI Lab62.9%2023-09-18
27Qwen-14BAlibaba61.3%2023-09-28
28GPT-3.5 Turbo (older v0613)OpenAI57.8%2023-06-13
29StarCoder 2 15BNVIDIA57.7%2024-02-20
30davinci-002OpenAI56.8%2022-03-15
31Palm 540BGoogle56.5%2022-04-04
32Mistral 7B v0.3Mistral AIopen54.4%32.12023-09-27
33Falcon-180BTII54.4%2023-09-06
34LLaMA-65BMetaopen54.4%2023-02-24
35Falcon 2 11BTII53.8%2024-05-09
36Baichuan2-13BBaichuan52.8%2023-09-06
37Qwen-7BAlibaba51.7%2023-09-28
38Gemma 7BGoogle46.4%2024-02-21
39Nemotron-4 15BNVIDIA46.0%2024-02-26
40Baichuan2-13BBaichuan45.7%2023-09-06
41Yi 6B01.AI44.9%2023-11-22
42LLaMA-33BMeta44.1%2023-02-24
43Llama 2-34BMeta42.2%2023-07-18
44davinci-002OpenAI41.5%2022-03-15
45INTELLECT-1Prime Intellect38.6%2024-11-29
46CodeQwen1.5-7BUnknown37.7%2024-04-15
47Llama 2-13BMetaopen ↗36.9%36.52023-07-18
48DeepSeek Coder 33BDeepSeek35.4%2023-11-02
49Qwen2.5-Coder-0.5BAlibaba34.5%2024-09-18
50MPT-30BMosaicML34.4%2023-06-22
51Falcon-40BTII33.8%2023-05-25
52PaLM 62BGoogle33.0%2022-04-04
53StarCoder 2 7BNVIDIA32.7%2024-02-20
54chatglm2-6bZhipu AI32.4%2023-06-24
55internlm-7bShanghai AI Lab31.2%2023-07-05
56vicuna-13b-v1.1LMSYS28.1%2023-04-12
57Baichuan 1-13BBaichuan26.8%2023-07-11
58Baichuan 2-7BBaichuan24.6%2023-09-20
59StarCoder 2 3BNVIDIA21.6%2024-02-22
60DeepSeek Coder 6.7BDeepSeek21.3%2023-11-02

Cite as: BenchLeader, “GSM8K leaderboard”, https://www.benchleader.com/benchmarks/gsm8k, data as of 19 Sept 2026.

GSM8K: questions

What does GSM8K measure?
Grade-school maths word problems. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads GSM8K?
Deepseek Coder v2 leads GSM8K with 94.5% as of 19 Sept 2026, ahead of Qwen2.5-Coder-14B at 94.2%.
How many models have GSM8K results?
74 model configurations have a GSM8K result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs GSM8K and how often is it updated?
GSM8K is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does GSM8K count toward the BenchLeader Index?
No. GSM8K is shown for reference but left out of the composite index.