BenchLeader

ARC (AI2 Challenge)

Grade-school science questions, challenge set. A 2018 benchmark, saturated for current models; kept for history.

As of 19 Sept 2026, DeepSeek V3 leads ARC (AI2 Challenge) on BenchLeader with 95.3%, ahead of Llama 3.1 405B at 95.3%, across 77 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Reasoning
Index weight
Reference only
Models
77
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Grade-school science questions, challenge set.

How it is scored

Accuracy, as compiled by Epoch AI from published results.

What to keep in mind

Saturated: frontier models score near the ceiling, so it separates only older or smaller models, and results are largely self-reported by developers.

77 of 77
#
1DeepSeek V3DeepSeekopen95.3%44.82024-12-26
2Llama 3.1 405BMetaopen95.3%43.52024-07-23
3Qwen2.5 72BAlibabaopen94.5%42.72024-09-19
4Deepseek v2DeepSeekopen ↗92.2%2024-05-07
5phi-3-medium 14BmediumMicrosoft91.6%2024-04-23
6Phi 3 SmallMicrosoft90.7%2024-04-23
7GPT 3.5 Turbo 1106OpenAI87.4%39.22023-11-06
8Mixtral 8x7BMistral AIopen87.3%32.12023-12-11
9Claude InstantAnthropic86.3%2023-08-09
10Stable Beluga 2Stability AI86.1%2023-07-20
11Claude InstantAnthropic85.7%
12davinci-002OpenAI85.2%2022-03-15
13Palm 540BGoogle85.2%2022-04-04
14Phi 3 MiniMicrosoftopen84.9%36.22024-04-23
15Qwen-14BAlibaba84.4%2023-09-28
16Llama 3-8BMetaopen82.8%35.62024-04-18
17internlm-20bShanghai AI Lab81.7%2023-09-18
18Mistral 7B v0.3Mistral AIopen78.6%32.12023-09-27
19Llama 2-70BMetaopen ↗78.3%34.92023-07-18
20Gemma 7BGoogle78.3%2024-02-21
21Phi-2Microsoft75.9%2023-12-12
22Qwen-7BAlibaba75.3%2023-09-28
23Qwen2.5-Coder 32B InstructAlibabaopen ↗70.5%42.62024-09-18
24LLaMA-65BMetaopen69.5%2023-02-24
25internlm-7bShanghai AI Lab69.5%2023-07-05
26PaLM 2-LGoogle69.2%2023-05-17
27Falcon-180BTII67.8%2023-09-06
28LLaMA-33BMeta67.5%2023-02-24
29Qwen2.5-Coder-14BAlibaba66.0%2024-09-18
30PaLM 2-MGoogle64.9%2023-05-17
31DeepSeek-Coder-V2 236BDeepSeek64.3%2024-06-17
32Falcon-40BTII61.9%2023-03-15
33chatglm2-6bZhipu AI61.0%2023-06-24
34Qwen2.5-Coder 7B InstructAlibabaopen ↗60.9%2024-09-18
35Llama 2-13BMetaopen ↗60.3%36.52023-07-18
36PaLM 2-SGoogle59.6%2023-05-17
37DeepSeek-Coder-V2-Lite-BaseDeepSeek57.3%2024-06-13
38Yi-9B01.AI55.6%2024-03-01
39Nemotron-4 15BNVIDIA55.5%2024-02-26
40INTELLECT-1Prime Intellect54.5%2024-11-29
41Llama 2-34BMeta54.5%2023-07-18
42InstructGPT 175BOpenAI53.2%2022-01-27
43Qwen-1_8BAlibaba53.2%2023-11-30
44Qwen2.5-Coder-3BAlibaba52.9%2024-09-18
45LLaMA-13BMeta52.7%2023-02-24
46PaLM 62BGoogle52.5%2022-04-04
47MPT-30BMosaicML50.6%2023-06-22
48Yi 6B01.AI50.3%2023-11-02
49Falcon-7BTII47.9%2023-04-24
50LLaMA-7BMeta47.6%2023-02-24
51StarCoder 2 15BNVIDIA47.2%2024-02-20
52Llama 2-7BMetaopen ↗45.9%34.22023-07-18
53Qwen2 5 Coder 1 5BAlibaba45.2%2024-09-18
54Phi-1.5Microsoft44.4%2023-09-11
55vicuna-13b-v1.1LMSYS43.2%2023-04-12
56MPT-7BMosaicML42.6%2023-05-05
57DeepSeek Coder 33BDeepSeek42.2%2023-11-02
58Gemma 2BGoogle42.1%2024-02-21
59XGen-7BSalesforce41.2%2023-06-27
60GPT-NeoX-20BEleutherAI41.1%2022-04-07

Cite as: BenchLeader, “ARC (AI2 Challenge) leaderboard”, https://www.benchleader.com/benchmarks/arc_ai2, data as of 19 Sept 2026.

ARC (AI2 Challenge): questions

What does ARC (AI2 Challenge) measure?
Grade-school science questions, challenge set. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads ARC (AI2 Challenge)?
DeepSeek V3 leads ARC (AI2 Challenge) with 95.3% as of 19 Sept 2026, ahead of Llama 3.1 405B at 95.3%.
How many models have ARC (AI2 Challenge) results?
77 model configurations have a ARC (AI2 Challenge) result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs ARC (AI2 Challenge) and how often is it updated?
ARC (AI2 Challenge) is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does ARC (AI2 Challenge) count toward the BenchLeader Index?
No. ARC (AI2 Challenge) is shown for reference but left out of the composite index.