BenchLeader

Mystery Game Puzzles

Deduction puzzles in the style of mystery games. Run by Epoch AI.

As of 19 Sept 2026, GPT-6 Astra leads Mystery Game Puzzles on BenchLeader with 84.0%, ahead of Claude Opus 5 at 59.0%, across 128 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Reasoning
Index weight
Reference only
Models
128
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Deduction puzzles with clues and suspects; the model must reason to the unique solution.

How it is scored

Percent solved, run by Epoch AI.

What to keep in mind

Puzzle styles are specific; models can be tuned to them.

128 of 128
#
1GPT-6 AstramaxOpenAI84.0%71.82026-09-03
2Claude Opus 5maxAnthropic59.0%69.92026-07-24
3Claude Fable 5.1maxAnthropic58.0%69.72026-09-01
4GPT-5.6 SolmaxOpenAI58.0%68.82026-07-09
5GPT-5.5xhighOpenAI56.0%67.62026-04-23
6GPT-5.5highOpenAI52.0%67.02026-04-23
7Claude Fable 5maxAnthropic52.0%65.82026-06-09
8Gemini 3.8 FlashhighGoogle47.0%64.52026-09-02
9DeepSeek V4 Pro 0813maxDeepSeekopen ↗43.0%56.02026-08-13
10Qwen3 8xhighAlibaba38.0%57.72026-08-02
11Claude Opus 5Anthropic37.0%66.72026-07-24
12GPT-5.4xhighOpenAI37.0%65.42026-03-05
13Gemini 3.7 FlashhighGoogle37.0%64.62026-08-13
14Claude Opus 4.8maxAnthropic36.0%64.22026-05-28
15GPT-5.6 TerramaxOpenAI35.0%65.02026-07-09
16Claude Sonnet 5maxAnthropic35.0%60.52026-06-30
17Grok 4.6xhighxAI34.0%64.62026-08-12
18Gemini 3.1 ProhighGoogle34.0%60.72026-02-19
19DeepSeek V4 Flash 0731maxDeepSeekopen ↗34.0%54.42026-07-31
20GLM-5.3maxZhipu AIopen ↗33.0%65.82026-08-14
21GPT-5.6 Solno reasoningOpenAI33.0%55.22026-07-09
22Gemini 3.5 FlashhighGoogle32.0%63.62026-05-19
23Qwen3 7maxAlibaba32.0%61.02026-05-19
24Gemini 3.1 PromediumGoogle32.0%2026-02-19
25Claude Opus 4.8xhighAnthropic31.0%2026-05-28
26Gemini 3.6 FlashhighGoogle30.0%61.72026-07-21
27o3highOpenAI29.0%53.02025-04-16
28Gemini 3.1 ProlowGoogle29.0%2026-02-19
29Claude Opus 4.7maxAnthropic28.0%64.32026-04-16
30GPT-5.5lowOpenAI28.0%61.22026-04-23
31GPT-5.4mediumOpenAI28.0%55.92026-03-05
32Gemini 3.5 FlashlowGoogle28.0%2026-05-19
33Kimi K3maxMoonshot AIopen ↗26.0%67.22026-07-16
34GPT-5.6 SollowOpenAI26.0%63.72026-07-09
35Gemini 3 FlashlowGoogle26.0%2025-12-17
36Qwen3 7no reasoningAlibaba26.0%2026-05-19
37Claude Opus 4.6maxAnthropic25.0%59.22026-02-05
38Gemini 3 FlashminimalGoogle25.0%54.92025-12-17
39Gemini 3.6 FlashminimalGoogle25.0%2026-07-21
40GPT-5highOpenAI23.0%58.62025-08-07
41GPT-5.2highOpenAI23.0%55.72025-12-11
42o3mediumOpenAI23.0%53.72025-04-16
43Gemini 3.6 FlashlowGoogle23.0%2026-07-21
44GPT-5.2mediumOpenAI22.0%58.72025-12-11
45Claude Opus 4.5Anthropic22.0%58.22025-11-24
46Qwen3.6 35B-A3Bno reasoningAlibaba22.0%43.62026-04-14
47GPT-5.6 LunamaxOpenAI21.0%60.22026-07-09
48Claude Opus 4.1Anthropic21.0%51.92025-08-05
49Gemini 3 FlashhighGoogle20.0%58.72025-12-17
50Qwen3.5 FlashAlibaba20.0%52.52026-02-25
51Nemotron 3 UltraNVIDIA20.0%48.32026-06-04
52GPT-5.6 Lunano reasoningOpenAI20.0%46.02026-07-09
53GPT-5.6 TerralowOpenAI19.0%57.82026-07-09
54Gemini 3.5 Flash LitelowGoogle19.0%2026-07-21
55GLM-5.2lowZhipu AIopen ↗19.0%2026-06-16
56GPT-5.1lowOpenAI19.0%2025-11-13
57Qwen3 6no reasoningAlibaba19.0%2026-04-20
58Kimi K2.6Moonshot AIopen ↗18.0%60.52026-04-20
59GPT-5.5no reasoningOpenAI18.0%54.02026-04-23
60Qwen3.5 397B-A17Bno reasoningAlibabaopen ↗18.0%51.72026-02-13

Cite as: BenchLeader, “Mystery Game Puzzles leaderboard”, https://www.benchleader.com/benchmarks/mystery_game_puzzles, data as of 19 Sept 2026.

Mystery Game Puzzles: questions

What does Mystery Game Puzzles measure?
Deduction puzzles with clues and suspects; the model must reason to the unique solution. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads Mystery Game Puzzles?
GPT-6 Astra leads Mystery Game Puzzles with 84.0% as of 19 Sept 2026, ahead of Claude Opus 5 at 59.0%.
How many models have Mystery Game Puzzles results?
128 model configurations have a Mystery Game Puzzles result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs Mystery Game Puzzles and how often is it updated?
Mystery Game Puzzles is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does Mystery Game Puzzles count toward the BenchLeader Index?
No. Mystery Game Puzzles is shown for reference but left out of the composite index.