BenchLeader
AnthropicReasoning modelAuto-detected

Claude Sonnet 4.6

Best configuration ranks #89 of 610 on the BenchLeader Index at 60.3 ±3.2. Last measured 8 Sept 2026. Released 17 Feb 2026.

Blended price
$6.00/M
$3.00 in · $15.00 out
Output speed
43 tok/s
First answer
2.12 s
first token 1.47 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index60
  2. Reasoning65
  3. Coding57
  4. Agents & tools55
  5. Maths63
  6. Knowledge62
  7. Instruction following57
  8. Human preference62
  9. Multimodal61
  10. Long context67
  11. Composite67

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
low55.6#17443 tok/s1.77 s$0.0057Agents & tools 71 · Composite 58 · Instruction following 44 · Knowledge 63 · Long context 61 · Multimodal 54
medium51.8#26643 tok/s2.12 s$0.0057Agents & tools 38 · Coding 61 · Composite 45 · Maths 60 · Reasoning 61
high52.3#25243 tok/s2.12 s$0.0057Agents & tools 49 · Knowledge 46 · Maths 57 · Reasoning 61
max49.7#30243 tok/s2.12 s$0.0057Agents & tools 36 · Coding 53 · Knowledge 44 · Maths 55 · Reasoning 60
defaultbest60.3#8943 tok/s2.12 s$0.0057Agents & tools 55 · Coding 57 · Composite 67 · Human preference 62 · Instruction following 57 · Knowledge 62 · Long context 67 · Maths 63 · Multimodal 61 · Reasoning 65

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarklowmediumhighmaxdefaultSourceTrend
GPQA Diamond83.3%#9183.3%#9178.8%#12087.4%#61Epoch AI Benchmarking Hub
LMArena Hard Prompts1504#19LMArena
LiveBench Reasoningnot in index84.8%#30LiveBench
GPQA Diamond (AA)not in index79.7%#19087.5%#96Artificial Analysis
Humanity's Last Exam (AA)not in index11.2%#23333.6%#89Artificial Analysis
GPQA Diamond (Vals)not in index85.6%#44Vals AI
ARC-AGI-186.5%#6786.0%#70ARC Prize
ARC-AGI-260.4%#5658.3%#62ARC Prize

Coding

BenchmarklowmediumhighmaxdefaultSourceTrend
SWE-bench Verified (Epoch)75.2%#14Epoch AI Benchmarking Hub
SciCode46.8%#79SciCode
WeirdML66.1%#37WeirdML
FrontierCode24.3%#23Cognition
LMArena Coding1528#15LMArena
LMArena WebDev1521#32LMArena
LiveBench Codingnot in index79.3%#16LiveBench
SciCode (AA)not in index50.1%#68Artificial Analysis
LiveCodeBench82.1%#55Vals AI
SWE-bench (Vals)not in index77.4%#34Vals AI

Agents & tools

BenchmarklowmediumhighmaxdefaultSourceTrend
Terminal-Bench53.4%#22Terminal-Bench
OSWorld-Verified 2.09.3%#108.3%#11OSWorld
APEX-Agents23.7%#32Mercor
LMArena Agent-1#27LMArena
LiveBench Agentic Codingnot in index42.6%#45LiveBench
Terminal-Bench Hard42.4%#4453.0%#17Artificial Analysis
τ²-Bench Telecom (AA)not in index79.0%#12579.5%#122Artificial Analysis
Terminal-Bench 2.1 (Vals)57.3%#34Vals AI
MCP Atlas69.5%#20Scale AI SEAL

Maths

BenchmarklowmediumhighmaxdefaultSourceTrend
OTIS Mock AIME82.2%#9375.6%#11171.1%#12085.8%#80Epoch AI Benchmarking Hub
ProofBench45.0%#22Vals AI
LiveBench Mathematicsnot in index87.0%#35LiveBench
AIME (Vals)92.3%#20Vals AI

Knowledge

BenchmarklowmediumhighmaxdefaultSourceTrend
SimpleQA Verified35.5%#4132.8%#51Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index78.0%#21LiveBench
AA-Omniscience-2.1#10412.2#60Artificial Analysis
MMLU-Pro87.3%#29Vals AI
LegalBench82.1%#67Vals AI
CorpFin65.3%#29Vals AI
TaxEval77.1%#4Vals AI
MedQA92.1%#36Vals AI

Instruction following

BenchmarklowmediumhighmaxdefaultSourceTrend
LiveBench Languagenot in index76.1%#36LiveBench
IFBench42.4%#23356.6%#139Artificial Analysis

Human preference

BenchmarklowmediumhighmaxdefaultSourceTrend
LMArena Text1472#35LMArena
EQ-Bench 41207#15EQ-Bench

Multimodal

BenchmarklowmediumhighmaxdefaultSourceTrend
LMArena Vision1282#25LMArena
MMMU-Pro69.2%#12473.3%#96Artificial Analysis

Long context

BenchmarklowmediumhighmaxdefaultSourceTrend
AA-LCR69.3%#17080.0%#51Artificial Analysis

Composite

BenchmarklowmediumhighmaxdefaultSourceTrend
Epoch Capabilities Indexnot in index152.3#34Epoch AI Benchmarking Hub
LiveBench73.0%#38LiveBench
AA Intelligence Index23.3#13330.4#77Artificial Analysis
Vals Indexnot in index50.6#28Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Amazon Bedrock (EU)69 tok/s1.17 s$3.30$16.501M
Claude Platform on AWS38 tok/s1.21 s$3.00$15.001M
Google Vertex (Global)37 tok/s1.95 s$3.00$15.001M
Azure36 tok/s1.97 s$3.00$15.001M
Anthropic32 tok/s1.73 s$3.00$15.001M
Amazon Bedrock (Global)29 tok/s1.04 s$3.00$15.001M

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$4.54$9.08$13.61$18.15Mar 26Apr 26May 26Jun 26Jul 26Aug 26
input outputnow $3.30 in · $16.50 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.300 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0057$0.00499.0 s
Summarise a 30-page report12,000 / 600$0.045$0.02115.9 s
Code edit6,000 / 1,500$0.041$0.02836.6 s
Agentic coding session60,000 / 4,000$0.240$0.1181.6 min
Structured extraction2,000 / 200$0.0090$0.00506.7 s

See also

Data as of 9 Sept 2026. Compare these configurations.