BenchLeader
AnthropicReasoning modelAuto-detected

Claude Fable 5

Best configuration ranks #2 of 610 on the BenchLeader Index at 70.4 ±3.2. Last measured 8 Sept 2026. Released 9 Jun 2026.

Blended price
$20.00/M
$10.00 in · $50.00 out
Output speed
63 tok/s
First answer
92 s
first token 5.42 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index70
  2. Reasoning76
  3. Coding71
  4. Agents & tools74
  5. Knowledge71
  6. Instruction following63
  7. Human preference71
  8. Multimodal70
  9. Long context68
  10. Composite92

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
low63 tok/s92 s$0.019Maths 67 · Reasoning 63
medium63 tok/s92 s$0.019Reasoning 69
high63.9#4063 tok/s92 s$0.019Agents & tools 68 · Coding 76 · Maths 68 · Reasoning 67
xhigh63 tok/s92 s$0.019Knowledge 76 · Reasoning 72
max65.8#2363 tok/s92 s$0.019Agents & tools 51 · Coding 75 · Composite 76 · Maths 76 · Reasoning 71
thinking63 tok/s92 s$0.019Reasoning 82
defaultbest70.4#263 tok/s92 s$0.019Agents & tools 74 · Coding 71 · Composite 92 · Human preference 71 · Instruction following 63 · Knowledge 71 · Long context 68 · Multimodal 70 · Reasoning 76

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarklowmediumhighxhighmaxthinkingdefaultSourceTrend
GPQA Diamond78.8%#12083.3%#9185.9%#72Epoch AI Benchmarking Hub
SimpleBench81.9%#1SimpleBench
LMArena Hard Prompts1533#1LMArena
LiveBench Reasoningnot in index89.7%#9LiveBench
GPQA Diamond (AA)not in index92.6%#32Artificial Analysis
Humanity's Last Exam (AA)not in index55.5%#4Artificial Analysis
GPQA Diamond (Vals)not in index93.2%#10Vals AI
EnigmaEval39.3%#1Scale AI SEAL
Kagi LLM Benchmark91.4%#188.8%#2Kagi LLM Benchmark
ARC-AGI-190.5%#5092.5%#3595.5%#2098.5%#198.5%#1ARC Prize
ARC-AGI-276.8%#2982.5%#2687.5%#1488.3%#1289.2%#10ARC Prize

Coding

BenchmarklowmediumhighxhighmaxthinkingdefaultSourceTrend
SciCode60.2%#2SciCode
WeirdML87.8%#991.9%#4WeirdML
FrontierCode53.5%#1Cognition
LMArena Coding1551#3LMArena
LMArena WebDev1628#9LMArena
LiveBench Codingnot in index86.0%#2LiveBench
SciCode (AA)not in index61.0%#2Artificial Analysis
LiveCodeBench89.8%#2Vals AI
IOI72.3%#5Vals AI
SWE-bench (Vals)not in index95.0%#7Vals AI

Agents & tools

BenchmarklowmediumhighxhighmaxthinkingdefaultSourceTrend
Terminal-Bench44.5%#32Terminal-Bench
Remote Labor Index16.1%#1Scale AI / CAIS
APEX-Agents45.0%#3Mercor
LMArena Agent9.2#5LMArena
LiveBench Agentic Codingnot in index62.2%#6LiveBench
Terminal-Bench Hard62.9%#2Artificial Analysis
τ²-Bench Telecom (AA)not in index98.5%#4Artificial Analysis
Terminal-Bench 2.1 (Vals)80.5%#7Vals AI
MCP Atlas83.3%#5Scale AI SEAL
HiL-Bench56.3%#3Scale AI SEAL

Maths

BenchmarklowmediumhighxhighmaxthinkingdefaultSourceTrend
FrontierMath Tiers 1–387.0%#5Epoch AI Benchmarking Hub
FrontierMath Tier 490.2%#5Epoch AI Benchmarking Hub
OTIS Mock AIME97.8%#17100.0%#199.7%#8Epoch AI Benchmarking Hub
ProofBench95.0%#4Vals AI
LiveBench Mathematicsnot in index96.0%#4LiveBench

Knowledge

BenchmarklowmediumhighxhighmaxthinkingdefaultSourceTrend
SimpleQA Verified70.7%#4Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index80.5%#3LiveBench
AA-Omniscience43.3#5Artificial Analysis
MMLU-Pro91.5%#3Vals AI
LegalBench88.6%#1Vals AI
CorpFin71.8%#2Vals AI
TaxEval76.9%#5Vals AI
PRBench Finance53.9%#2Scale AI SEAL
PRBench Legal52.6%#2Scale AI SEAL

Instruction following

BenchmarklowmediumhighxhighmaxthinkingdefaultSourceTrend
LiveBench Languagenot in index90.7%#1LiveBench
IFBench63.5%#112Artificial Analysis

Human preference

BenchmarklowmediumhighxhighmaxthinkingdefaultSourceTrend
LMArena Text1507#1LMArena
EQ-Bench 41340#2EQ-Bench

Multimodal

BenchmarklowmediumhighxhighmaxthinkingdefaultSourceTrend
LMArena Vision1330#1LMArena

Long context

BenchmarklowmediumhighxhighmaxthinkingdefaultSourceTrend
AA-LCR82.3%#21Artificial Analysis

Composite

BenchmarklowmediumhighxhighmaxthinkingdefaultSourceTrend
Epoch Capabilities Indexnot in index163.4#3Epoch AI Benchmarking Hub
LiveBench83.0%#2LiveBench
AA Intelligence Index49.7#8Artificial Analysis
Vals Indexnot in index66.0#4Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Google Vertex64 tok/s3.86 s$10.00$50.001M
Anthropic44 tok/s5.45 s$10.00$50.001M
Azure (BYOK Only)35 tok/s8.56 s$10.00$50.001M
Claude Platform on AWS22 tok/s5.40 s$10.00$50.001M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $1.00 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.019$0.0161.6 min
Summarise a 30-page report12,000 / 600$0.150$0.0691.7 min
Code edit6,000 / 1,500$0.135$0.0951.9 min
Agentic coding session60,000 / 4,000$0.800$0.3952.6 min
Structured extraction2,000 / 200$0.030$0.0171.6 min

See also

Data as of 9 Sept 2026. Compare these configurations.