DTBench
Decision-theory dilemmas scored against expert answers.
As of 19 Sept 2026, Claude Fable 5 leads DTBench on BenchLeader with 98.4%, ahead of Claude Opus 5 at 97.6%, across 154 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Reasoning
- Index weight
- Reference only
- Models
- 154
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Decision problems where the model's answer preference is compared with decision-theoretic expert answers.
How it is scored
Accuracy, as published by DTBench.
What to keep in mind
Philosophy-of-decision framing; a large model set.
- 1Claude Fable 5 (max)98.4%
- 2Claude Opus 5 (max)97.6%
- 3Grok 4.6 (xhigh)97.3%
- 4Gemini 3.7 Flash (high)96.8%
- 5Grok 4.5 (high)96.5%
- 6GPT-5.5 (xhigh)96.0%
- 7GPT-5.5 Pro (xhigh)96.0%
- 8GPT-5.6 Sol96.0%
- 9GPT-5.6 Sol (max)95.5%
- 10Gemini 3.6 Flash (high)95.5%
- 11Gemini 3.1 Pro (high)95.5%
- 12Claude Opus 4.8 (max)94.9%
- 13Muse Spark 1.2 (xhigh)94.7%
- 14Gemini 3.5 Flash (high)94.7%
- 15Claude Opus 4.7 (max)94.7%
154 of 154
| # | ||||
|---|---|---|---|---|
| 1 | 98.4% | 65.8 | 2026-06-09 | |
| 2 | 97.6% | 69.9 | 2026-07-24 | |
| 3 | 97.3% | 64.6 | 2026-08-12 | |
| 4 | 96.8% | 64.6 | 2026-08-13 | |
| 5 | 96.5% | 61.0 | 2026-07-08 | |
| 6 | 96.0% | 67.6 | 2026-04-23 | |
| 7 | 96.0% | 67.1 | 2026-04-23 | |
| 8 | 96.0% | – | 2026-07-09 | |
| 9 | 95.5% | 68.8 | 2026-07-09 | |
| 10 | 95.5% | 61.7 | 2026-07-21 | |
| 11 | 95.5% | 60.7 | 2026-02-19 | |
| 12 | 94.9% | 64.2 | 2026-05-28 | |
| 13 | 94.7% | 64.1 | 2026-08-05 | |
| 14 | 94.7% | 63.6 | 2026-05-19 | |
| 15 | 94.7% | 64.3 | 2026-04-16 | |
| 16 | 94.4% | 65.4 | 2026-03-05 | |
| 17 | 94.4% | – | 2026-07-09 | |
| 18 | 93.6% | 63.9 | 2026-06-16 | |
| 19 | 92.5% | 65.0 | 2026-07-09 | |
| 20 | 92.5% | 60.5 | 2026-06-30 | |
| 21 | 92.3% | 61.0 | 2026-05-19 | |
| 22 | 92.0% | 57.7 | 2026-08-02 | |
| 23 | 91.2% | 67.2 | 2026-07-16 | |
| 24 | 91.2% | 59.2 | 2026-02-05 | |
| 25 | 90.9% | 62.0 | 2025-12-11 | |
| 26 | 90.9% | 60.5 | 2026-04-20 | |
| 27 | 90.9% | 54.4 | 2026-07-31 | |
| 28 | 90.7% | 64.2 | 2026-04-24 | |
| 29 | 90.7% | 58.6 | 2025-08-07 | |
| 30 | 90.7% | 58.2 | 2026-04-17 | |
| 31 | 90.1% | 58.7 | 2025-11-13 | |
| 32 | 90.1% | 48.3 | 2026-06-04 | |
| 33 | 90.1% | – | 2026-02-17 | |
| 34 | 89.9% | 57.7 | 2025-11-24 | |
| 35 | 89.9% | 57.6 | 2026-02-17 | |
| 36 | 89.1% | 60.2 | 2026-07-09 | |
| 37 | 89.1% | 58.7 | 2025-12-17 | |
| 38 | 87.7% | 56.1 | 2025-09-29 | |
| 39 | 87.7% | 56.0 | 2025-11-19 | |
| 40 | 87.5% | 56.9 | 2026-02-13 | |
| 41 | 87.2% | 63.3 | 2026-04-20 | |
| 42 | 86.9% | 54.5 | 2025-06-10 | |
| 43 | 86.4% | 62.1 | 2026-04-24 | |
| 44 | 84.8% | 53.0 | 2025-04-16 | |
| 45 | 84.5% | 59.2 | – | |
| 46 | 84.3% | 48.1 | 2026-02-24 | |
| 47 | 84.0% | 59.9 | 2026-06-02 | |
| 48 | 83.5% | 44.3 | 2026-07-21 | |
| 49 | 83.2% | 54.3 | 2025-09-29 | |
| 50 | 82.9% | 52.5 | 2026-02-25 | |
| 51 | 82.7% | 54.3 | 2025-09-19 | |
| 52 | 82.7% | 53.9 | 2025-08-21 | |
| 53 | 82.7% | 52.7 | 2026-04-02 | |
| 54 | 82.4% | 54.5 | 2025-06-17 | |
| 55 | 82.4% | 53.0 | 2026-02-24 | |
| 56 | 82.1% | 53.5 | 2025-09-24 | |
| 57 | 81.9% | 58.5 | 2026-03-31 | |
| 58 | 81.6% | 52.9 | 2025-05-22 | |
| 59 | 81.3% | 54.3 | 2025-09-22 | |
| 60 | 81.1% | 49.1 | 2025-04-28 |
Cite as: BenchLeader, “DTBench leaderboard”, https://www.benchleader.com/benchmarks/dtbench, data as of 19 Sept 2026.
DTBench: questions
- What does DTBench measure?
- Decision problems where the model's answer preference is compared with decision-theoretic expert answers. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads DTBench?
- Claude Fable 5 leads DTBench with 98.4% as of 19 Sept 2026, ahead of Claude Opus 5 at 97.6%.
- How many models have DTBench results?
- 154 model configurations have a DTBench result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs DTBench and how often is it updated?
- DTBench is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does DTBench count toward the BenchLeader Index?
- No. DTBench is shown for reference but left out of the composite index.