BenchLeader

Chess Puzzles

Solving chess puzzles from a board description. Run by Epoch AI.

As of 19 Sept 2026, GPT-6 Astra leads Chess Puzzles on BenchLeader with 72.0%, ahead of GPT-5.6 Sol at 64.0%, across 217 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Reasoning
Index weight
Reference only
Models
217
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Chess puzzles given as a position; the model must find the winning move sequence.

How it is scored

Percent of puzzles solved, run by Epoch AI.

What to keep in mind

A narrow reasoning task with a large model set; good for spotting search-like reasoning.

217 of 217
#
1GPT-6 AstramaxOpenAI72.0%71.82026-09-03
2GPT-5.6 SolOpenAI64.0%2026-07-09
3GPT-5.5 ProxhighOpenAI64.0%2026-04-23
4Gemini 3.8 FlashhighGoogle61.0%64.52026-09-02
5GPT-5.4 ProxhighOpenAI58.6%63.32026-03-05
6GPT-5.6 SolmaxOpenAI55.0%68.82026-07-09
7Gemini 3.1 ProGoogle55.0%63.92026-02-19
8GPT-5.6 TerramaxOpenAI54.0%65.02026-07-09
9GPT-5.5xhighOpenAI54.0%2026-04-23
10Gemini 3.5 FlashhighGoogle50.0%63.62026-05-19
11GPT-5.2xhighOpenAI49.0%62.02025-12-11
12Gemini 3.1 ProhighGoogle49.0%60.72026-02-19
13Claude Fable 5.1maxAnthropic47.0%69.72026-09-01
14Gemini 3.7 FlashhighGoogle47.0%64.62026-08-13
15DeepSeek V4 Pro 0813maxDeepSeekopen ↗47.0%56.02026-08-13
16Gemini 3.5 FlashlowGoogle45.0%2026-05-19
17GPT-5.4xhighOpenAI44.0%65.42026-03-05
18Gemini 3.5 FlashminimalGoogle43.0%57.42026-05-19
19Gemini 3.6 FlashlowGoogle43.0%2026-07-21
20Claude Opus 5maxAnthropic42.0%69.92026-07-24
21Claude Fable 5maxAnthropic41.0%65.82026-06-09
22Claude Fable 5highAnthropic41.0%64.02026-06-09
23Grok 4.6highxAI40.0%64.52026-08-12
24Gemini 3.6 FlashhighGoogle40.0%61.72026-07-21
25GPT-5.6 LunamaxOpenAI40.0%60.22026-07-09
26Gemini 3 FlashhighGoogle40.0%58.72025-12-17
27GPT-5.2mediumOpenAI40.0%58.72025-12-11
28Qwen3.8 Max (0902)xhighAlibaba40.0%57.12026-09-01
29GPT-5.2highOpenAI40.0%55.72025-12-11
30Kimi K3maxMoonshot AIopen ↗39.0%67.22026-07-16
31GPT-5.4highOpenAI38.0%59.02026-03-05
32Gemini 3 FlashGoogle38.0%57.72025-12-17
33GPT-5.4mediumOpenAI38.0%55.92026-03-05
34o3mediumOpenAI38.0%53.72025-04-16
35GPT-5highOpenAI37.0%58.62025-08-07
36Grok 4.5highxAI36.0%61.02026-07-08
37Muse Spark 1.3xhighMeta35.0%68.02026-09-02
38Claude Sonnet 5xhighAnthropic35.0%59.82026-06-30
39Gemini 3.6 FlashminimalGoogle35.0%2026-07-21
40Claude Opus 4.8maxAnthropic34.0%64.22026-05-28
41o3highOpenAI34.0%53.02025-04-16
42Claude Opus 5Anthropic33.0%66.72026-07-24
43DeepSeek V4 Flash 0731maxDeepSeekopen ↗33.0%54.42026-07-31
44GPT-5.1highOpenAI32.0%58.72025-11-13
45Grok 4.6xhighxAI31.0%64.62026-08-12
46Gemini 3 ProGoogle31.0%61.12025-11-18
47Claude Opus 4.7xhighAnthropic30.0%57.32026-04-16
48GPT-5 minihighOpenAI30.0%55.32025-08-07
49GPT-5.4 nanohighOpenAI30.0%49.12026-03-17
50GPT-5mediumOpenAI29.0%58.52025-08-07
51Qwen3 8xhighAlibaba29.0%57.72026-08-02
52Claude Opus 4.8lowAnthropic29.0%2026-05-28
53Grok 4xAI28.0%58.52025-07-09
54Claude Fable 5lowAnthropic28.0%2026-06-09
55GPT-5.6 SollowOpenAI27.0%63.72026-07-09
56GPT-5 nanohighOpenAI27.0%48.22025-08-07
57o3lowOpenAI27.0%47.22025-04-16
58GPT-5.5lowOpenAI26.0%61.22026-04-23
59Kimi K2.6Moonshot AIopen ↗26.0%60.52026-04-20
60o4-minihighOpenAI26.0%50.82025-04-16

Cite as: BenchLeader, “Chess Puzzles leaderboard”, https://www.benchleader.com/benchmarks/chess_puzzles, data as of 19 Sept 2026.

Chess Puzzles: questions

What does Chess Puzzles measure?
Chess puzzles given as a position; the model must find the winning move sequence. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads Chess Puzzles?
GPT-6 Astra leads Chess Puzzles with 72.0% as of 19 Sept 2026, ahead of GPT-5.6 Sol at 64.0%.
How many models have Chess Puzzles results?
217 model configurations have a Chess Puzzles result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs Chess Puzzles and how often is it updated?
Chess Puzzles is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does Chess Puzzles count toward the BenchLeader Index?
No. Chess Puzzles is shown for reference but left out of the composite index.