BenchLeader
AnthropicReasoning model

Claude Sonnet 5

Best configuration ranks #72 of 610 on the BenchLeader Index at 61.2 ±2.9. Last measured 5 Sept 2026. Released 30 Jun 2026.

Blended price
$4.00/M
$2.00 in · $10.00 out
Output speed
79 tok/s
First answer
193 s
first token 2.75 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index61
  2. Reasoning59
  3. Coding62
  4. Agents & tools61
  5. Knowledge64
  6. Human preference56
  7. Multimodal62
  8. Long context68
  9. Composite77

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning59 tok/s1.01 s$0.0038Composite 65 · Knowledge 64 · Long context 61 · Multimodal 56
low62 tok/s1.39 s$0.0038Composite 60 · Knowledge 60 · Long context 60
medium59 tok/s1.48 s$0.0038Composite 65 · Knowledge 61 · Long context 63
high60.5#8463 tok/s14 s$0.0038Agents & tools 62 · Coding 62 · Human preference 66 · Knowledge 62 · Long context 65 · Multimodal 63 · Reasoning 67
xhigh57.3#14366 tok/s36 s$0.0038Composite 55 · Knowledge 55 · Long context 65 · Maths 66 · Reasoning 65
max50.7#28179 tok/s193 s$0.0038Agents & tools 28 · Coding 63 · Knowledge 45 · Maths 61 · Reasoning 59
defaultbest61.2#7279 tok/s193 s$0.0038Agents & tools 61 · Coding 62 · Composite 77 · Human preference 56 · Knowledge 64 · Long context 68 · Multimodal 62 · Reasoning 59

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
GPQA Diamond90.5%#3580.3%#112Epoch AI Benchmarking Hub
SimpleBench60.6%#29SimpleBench
LMArena Hard Prompts1490#41LMArena
LiveBench Reasoningnot in index88.7%#14LiveBench
GPQA Diamond (AA)not in index80.0%#18891.1%#50Artificial Analysis
Humanity's Last Exam (AA)not in index19.1%#16821.9%#15130.0%#10435.7%#7539.0%#6041.3%#48Artificial Analysis
GPQA Diamond (Vals)not in index88.9%#31Vals AI

Coding

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
SciCode48.6%#6953.6%#37SciCode
WeirdML68.8%#32WeirdML
FrontierCode42.7%#11Cognition
GSO-Bench37.3%#6GSO-Bench
LMArena Coding1521#24LMArena
LMArena WebDev1538#28LMArena
LiveBench Codingnot in index80.7%#12LiveBench
SciCode (AA)not in index50.1%#6851.6%#5654.3%#3954.0%#4254.3%#39Artificial Analysis
LiveCodeBench82.4%#51Vals AI
SWE-bench (Vals)not in index79.6%#25Vals AI

Agents & tools

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
Terminal-Bench12.4%#61Terminal-Bench
APEX-Agents32.5%#23Mercor
LMArena Agent6.3#9LMArena
LiveBench Agentic Codingnot in index59.4%#13LiveBench
Terminal-Bench 2.1 (Vals)74.5%#14Vals AI

Maths

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
FrontierMath Tiers 1–365.6%#23Epoch AI Benchmarking Hub
FrontierMath Tier 429.3%#28Epoch AI Benchmarking Hub
OTIS Mock AIME94.7%#3880.0%#99Epoch AI Benchmarking Hub
ProofBench77.0%#7Vals AI
LiveBench Mathematicsnot in index92.9%#14LiveBench

Knowledge

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
SimpleQA Verified32.9%#5033.7%#47Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index71.7%#35LiveBench
AA-Omniscience-0.7#98-8.2#128-6.9#123-3.7#1103.2#8116.4#50Artificial Analysis
MMLU-Pro87.5%#26Vals AI
LegalBench83.9%#42Vals AI
CorpFin67.0%#14Vals AI
TaxEval75.6%#14Vals AI

Instruction following

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
LiveBench Languagenot in index75.0%#39LiveBench

Human preference

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
LMArena Text1462#48LMArena
EQ-Bench 41236#10EQ-Bench

Multimodal

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
LMArena Vision1278#29LMArena
MMMU-Pro71.9%#10477.3%#59Artificial Analysis

Long context

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
AA-LCR70.0%#16167.3%#18273.7%#12576.7%#9676.7%#9682.0%#25Artificial Analysis

Composite

Benchmarkno reasoninglowmediumhighxhighmaxdefaultSourceTrend
Epoch Capabilities Indexnot in index156.2#19Epoch AI Benchmarking Hub
LiveBench76.0%#24LiveBench
AA Intelligence Index28.9#8624.7#12028.4#8938.4#45Artificial Analysis
Vals Indexnot in index59.6#11Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Google Vertex (Europe)81 tok/s3.68 s$2.20$11.001M
Google Vertex (Global)69 tok/s3.52 s$2.00$10.001M
Amazon Bedrock (US)67 tok/s4.43 s$2.20$11.001M
Anthropic66 tok/s1.98 s$2.00$10.001M
Azure (US)64 tok/s0.95 s$2.00$10.001M
Azure62 tok/s1.19 s$2.00$10.001M
Claude Platform on AWS51 tok/s3.80 s$2.00$10.001M
Amazon Bedrock (Global)37 tok/s1.91 s$2.00$10.001M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.200 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0038$0.00333.3 min
Summarise a 30-page report12,000 / 600$0.030$0.0143.3 min
Code edit6,000 / 1,500$0.027$0.0193.5 min
Agentic coding session60,000 / 4,000$0.160$0.0794.1 min
Structured extraction2,000 / 200$0.0060$0.00333.3 min

See also

Data as of 9 Sept 2026. Compare these configurations.