BenchLeader

DTBench

Decision-theory dilemmas scored against expert answers.

As of 19 Sept 2026, Claude Fable 5 leads DTBench on BenchLeader with 98.4%, ahead of Claude Opus 5 at 97.6%, across 154 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Reasoning
Index weight
Reference only
Models
154
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Decision problems where the model's answer preference is compared with decision-theoretic expert answers.

How it is scored

Accuracy, as published by DTBench.

What to keep in mind

Philosophy-of-decision framing; a large model set.

154 of 154
#
1Claude Fable 5maxAnthropic98.4%65.82026-06-09
2Claude Opus 5maxAnthropic97.6%69.92026-07-24
3Grok 4.6xhighxAI97.3%64.62026-08-12
4Gemini 3.7 FlashhighGoogle96.8%64.62026-08-13
5Grok 4.5highxAI96.5%61.02026-07-08
6GPT-5.5xhighOpenAI96.0%67.62026-04-23
7GPT-5.5 ProxhighOpenAI96.0%67.12026-04-23
8GPT-5.6 SolOpenAI96.0%2026-07-09
9GPT-5.6 SolmaxOpenAI95.5%68.82026-07-09
10Gemini 3.6 FlashhighGoogle95.5%61.72026-07-21
11Gemini 3.1 ProhighGoogle95.5%60.72026-02-19
12Claude Opus 4.8maxAnthropic94.9%64.22026-05-28
13Muse Spark 1.2xhighMeta94.7%64.12026-08-05
14Gemini 3.5 FlashhighGoogle94.7%63.62026-05-19
15Claude Opus 4.7maxAnthropic94.7%64.32026-04-16
16GPT-5.4xhighOpenAI94.4%65.42026-03-05
17Muse Spark 1.1highMeta94.4%2026-07-09
18GLM-5.2maxZhipu AIopen ↗93.6%63.92026-06-16
19GPT-5.6 TerramaxOpenAI92.5%65.02026-07-09
20Claude Sonnet 5maxAnthropic92.5%60.52026-06-30
21Qwen3 7maxAlibaba92.3%61.02026-05-19
22Qwen3 8xhighAlibaba92.0%57.72026-08-02
23Kimi K3maxMoonshot AIopen ↗91.2%67.22026-07-16
24Claude Opus 4.6maxAnthropic91.2%59.22026-02-05
25GPT-5.2xhighOpenAI90.9%62.02025-12-11
26Kimi K2.6Moonshot AIopen ↗90.9%60.52026-04-20
27DeepSeek V4 Flash 0731maxDeepSeekopen ↗90.9%54.42026-07-31
28DeepSeek V4 PromaxDeepSeekopen ↗90.7%64.22026-04-24
29GPT-5highOpenAI90.7%58.62025-08-07
30Grok 4.3highxAI90.7%58.22026-04-17
31GPT-5.1highOpenAI90.1%58.72025-11-13
32Nemotron 3 UltraNVIDIA90.1%48.32026-06-04
33Grok 4.20xAI90.1%2026-02-17
34Claude Opus 4.5highAnthropic89.9%57.72025-11-24
35Claude Sonnet 4.6maxAnthropic89.9%57.62026-02-17
36GPT-5.6 LunamaxOpenAI89.1%60.22026-07-09
37Gemini 3 FlashhighGoogle89.1%58.72025-12-17
38DeepSeek V3.2thinkingDeepSeekopen87.7%56.12025-09-29
39Grok 4.1thinkingxAI87.7%56.02025-11-19
40Qwen3.5 397B-A17BAlibabaopen ↗87.5%56.92026-02-13
41Qwen3 6maxAlibaba87.2%63.32026-04-20
42o3-prohighOpenAI86.9%54.52025-06-10
43DeepSeek V4 FlashmaxDeepSeekopen ↗86.4%62.12026-04-24
44o3highOpenAI84.8%53.02025-04-16
45MiMo-V2.5-ProXiaomiopen ↗84.5%59.2
46Qwen3.5 122B-A10Bno reasoningAlibaba84.3%48.12026-02-24
47Qwen3.7 PlusAlibaba84.0%59.92026-06-02
48Gemini 3.5 Flash LitehighGoogle83.5%44.32026-07-21
49Claude Sonnet 4.5Anthropic83.2%54.32025-09-29
50Qwen3.5 FlashAlibaba82.9%52.52026-02-25
51Grok 4 FastxAI82.7%54.32025-09-19
52DeepSeek V3.1thinkingDeepSeekopen82.7%53.92025-08-21
53Gemma 4 31BGoogleopen ↗82.7%52.72026-04-02
54Gemini 2.5 ProGoogle82.4%54.52025-06-17
55Qwen3.5 27BAlibaba82.4%53.02026-02-24
56Qwen3 MaxmaxAlibaba82.1%53.52025-09-24
57Qwen3.6 PlusAlibabaopen81.9%58.52026-03-31
58Claude Opus 4Anthropic81.6%52.92025-05-22
59DeepSeek V3.1DeepSeekopen81.3%54.32025-09-22
60Qwen PlusAlibaba81.1%49.12025-04-28

Cite as: BenchLeader, “DTBench leaderboard”, https://www.benchleader.com/benchmarks/dtbench, data as of 19 Sept 2026.

DTBench: questions

What does DTBench measure?
Decision problems where the model's answer preference is compared with decision-theoretic expert answers. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads DTBench?
Claude Fable 5 leads DTBench with 98.4% as of 19 Sept 2026, ahead of Claude Opus 5 at 97.6%.
How many models have DTBench results?
154 model configurations have a DTBench result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs DTBench and how often is it updated?
DTBench is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does DTBench count toward the BenchLeader Index?
No. DTBench is shown for reference but left out of the composite index.