Chess Puzzles
Solving chess puzzles from a board description. Run by Epoch AI.
As of 19 Sept 2026, GPT-6 Astra leads Chess Puzzles on BenchLeader with 72.0%, ahead of GPT-5.6 Sol at 64.0%, across 217 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Reasoning
- Index weight
- Reference only
- Models
- 217
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Chess puzzles given as a position; the model must find the winning move sequence.
How it is scored
Percent of puzzles solved, run by Epoch AI.
What to keep in mind
A narrow reasoning task with a large model set; good for spotting search-like reasoning.
- 1GPT-6 Astra (max)72.0%
- 2GPT-5.6 Sol64.0%
- 3GPT-5.5 Pro (xhigh)64.0%
- 4Gemini 3.8 Flash (high)61.0%
- 5GPT-5.4 Pro (xhigh)58.6%
- 6GPT-5.6 Sol (max)55.0%
- 7Gemini 3.1 Pro55.0%
- 8GPT-5.6 Terra (max)54.0%
- 9GPT-5.5 (xhigh)54.0%
- 10Gemini 3.5 Flash (high)50.0%
- 11GPT-5.2 (xhigh)49.0%
- 12Gemini 3.1 Pro (high)49.0%
- 13Claude Fable 5.1 (max)47.0%
- 14Gemini 3.7 Flash (high)47.0%
- 15DeepSeek V4 Pro 0813 (max)47.0%
217 of 217
| # | ||||
|---|---|---|---|---|
| 1 | 72.0% | 71.8 | 2026-09-03 | |
| 2 | 64.0% | – | 2026-07-09 | |
| 3 | 64.0% | – | 2026-04-23 | |
| 4 | 61.0% | 64.5 | 2026-09-02 | |
| 5 | 58.6% | 63.3 | 2026-03-05 | |
| 6 | 55.0% | 68.8 | 2026-07-09 | |
| 7 | 55.0% | 63.9 | 2026-02-19 | |
| 8 | 54.0% | 65.0 | 2026-07-09 | |
| 9 | 54.0% | – | 2026-04-23 | |
| 10 | 50.0% | 63.6 | 2026-05-19 | |
| 11 | 49.0% | 62.0 | 2025-12-11 | |
| 12 | 49.0% | 60.7 | 2026-02-19 | |
| 13 | 47.0% | 69.7 | 2026-09-01 | |
| 14 | 47.0% | 64.6 | 2026-08-13 | |
| 15 | 47.0% | 56.0 | 2026-08-13 | |
| 16 | 45.0% | – | 2026-05-19 | |
| 17 | 44.0% | 65.4 | 2026-03-05 | |
| 18 | 43.0% | 57.4 | 2026-05-19 | |
| 19 | 43.0% | – | 2026-07-21 | |
| 20 | 42.0% | 69.9 | 2026-07-24 | |
| 21 | 41.0% | 65.8 | 2026-06-09 | |
| 22 | 41.0% | 64.0 | 2026-06-09 | |
| 23 | 40.0% | 64.5 | 2026-08-12 | |
| 24 | 40.0% | 61.7 | 2026-07-21 | |
| 25 | 40.0% | 60.2 | 2026-07-09 | |
| 26 | 40.0% | 58.7 | 2025-12-17 | |
| 27 | 40.0% | 58.7 | 2025-12-11 | |
| 28 | 40.0% | 57.1 | 2026-09-01 | |
| 29 | 40.0% | 55.7 | 2025-12-11 | |
| 30 | 39.0% | 67.2 | 2026-07-16 | |
| 31 | 38.0% | 59.0 | 2026-03-05 | |
| 32 | 38.0% | 57.7 | 2025-12-17 | |
| 33 | 38.0% | 55.9 | 2026-03-05 | |
| 34 | 38.0% | 53.7 | 2025-04-16 | |
| 35 | 37.0% | 58.6 | 2025-08-07 | |
| 36 | 36.0% | 61.0 | 2026-07-08 | |
| 37 | 35.0% | 68.0 | 2026-09-02 | |
| 38 | 35.0% | 59.8 | 2026-06-30 | |
| 39 | 35.0% | – | 2026-07-21 | |
| 40 | 34.0% | 64.2 | 2026-05-28 | |
| 41 | 34.0% | 53.0 | 2025-04-16 | |
| 42 | 33.0% | 66.7 | 2026-07-24 | |
| 43 | 33.0% | 54.4 | 2026-07-31 | |
| 44 | 32.0% | 58.7 | 2025-11-13 | |
| 45 | 31.0% | 64.6 | 2026-08-12 | |
| 46 | 31.0% | 61.1 | 2025-11-18 | |
| 47 | 30.0% | 57.3 | 2026-04-16 | |
| 48 | 30.0% | 55.3 | 2025-08-07 | |
| 49 | 30.0% | 49.1 | 2026-03-17 | |
| 50 | 29.0% | 58.5 | 2025-08-07 | |
| 51 | 29.0% | 57.7 | 2026-08-02 | |
| 52 | 29.0% | – | 2026-05-28 | |
| 53 | 28.0% | 58.5 | 2025-07-09 | |
| 54 | 28.0% | – | 2026-06-09 | |
| 55 | 27.0% | 63.7 | 2026-07-09 | |
| 56 | 27.0% | 48.2 | 2025-08-07 | |
| 57 | 27.0% | 47.2 | 2025-04-16 | |
| 58 | 26.0% | 61.2 | 2026-04-23 | |
| 59 | 26.0% | 60.5 | 2026-04-20 | |
| 60 | 26.0% | 50.8 | 2025-04-16 |
Cite as: BenchLeader, “Chess Puzzles leaderboard”, https://www.benchleader.com/benchmarks/chess_puzzles, data as of 19 Sept 2026.
Chess Puzzles: questions
- What does Chess Puzzles measure?
- Chess puzzles given as a position; the model must find the winning move sequence. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads Chess Puzzles?
- GPT-6 Astra leads Chess Puzzles with 72.0% as of 19 Sept 2026, ahead of GPT-5.6 Sol at 64.0%.
- How many models have Chess Puzzles results?
- 217 model configurations have a Chess Puzzles result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs Chess Puzzles and how often is it updated?
- Chess Puzzles is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does Chess Puzzles count toward the BenchLeader Index?
- No. Chess Puzzles is shown for reference but left out of the composite index.