BenchLeader

Claude Sonnet 5.5

Reasoning effort

Claude Sonnet 5.5 is an Anthropic proprietary reasoning model, released 28 Sept 2026. Its best configuration (max reasoning effort) ranks #21 of 760 on the BenchLeader Index at 67.6 ±5.6, in the top 25. The ± is the point: 87 other configurations score within that range, so they and this one cannot be told apart on quality alone — price and speed are what separate them. It scores highest in reasoning (81) and lowest in knowledge (65). At $4.00 per million tokens blended it is among the most expensive ranked models. Output speed of 141 tokens per second puts it in the fastest quarter, with a first answer in 487.1 s. It has been measured at 6 reasoning-effort settings; this summary describes the best-scoring one, and the tabs above switch between them. Last measured 8 Oct 2026.

Blended price
$4.00/M
$2.00 in · $10.00 out
Output speed
141 tok/s
measured by Artificial Analysis
First answer
487 s
first token 2.69 s
Context
1M
Full answer
491 s
median, reasoning included
Cost per run
$5.46
one full Intelligence Index run
Released
28 Sept 2026
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index68
  2. Reasoning81
  3. Coding70
  4. Agents & tools70
  5. Maths73
  6. Knowledge65
  7. Long context67
  8. Composite73

Versions

Anthropic has shipped 9 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
Claude Sonnet 5.5maxthis page28 Sept 202667.6#21
Claude Sonnet 5high30 Jun 202660.5#102
Claude Sonnet 4.617 Feb 202658.6#136
Claude Sonnet 4.5high29 Sept 202558.3#144
Claude Sonnet 4thinking22 May 202553.3#275
Claude 3.7 Sonnetthinking24 Feb 202551.1#336
Claude 3.5 Sonnet22 Oct 202446.4#468
Claude 3 Sonnet29 Feb 202438.2#692
Claude 37 Sonnetthinking–––

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Not stated” collects results from publishers that did not say which setting they used; for a reasoning model that is usually its thinking mode, but we do not assume it. Pick a setting here or at the top of the page to see its price, speed and category scores.

EffortIndexRankSpeedFirst answerChat reply costCategories
low58.4#143100 tok/s0.88 s$0.0038Coding 54 · Composite 70 · Knowledge 70 · Long context 63 · Reasoning 60
medium59.1#12499 tok/s1.87 s$0.0038Agents & tools 44 · Coding 59 · Composite 76 · Knowledge 71 · Long context 63 · Reasoning 69
high66.1#32101 tok/s12 s$0.0038Coding 70 · Composite 83 · Knowledge 71 · Long context 64 · Reasoning 82
xhigh67.1#26104 tok/s30 s$0.0038Coding 72 · Composite 73 · Human preference 67 · Knowledge 72 · Long context 65 · Maths 70 · Multimodal 63 · Reasoning 85
maxbest67.6#21141 tok/s487 s$0.0038Agents & tools 70 · Coding 70 · Composite 73 · Knowledge 65 · Long context 67 · Maths 73 · Reasoning 81
not stated––97 tok/s2.69 s$0.0038Agents & tools 69 · Coding 65

What more thinking costs

Turning the effort up buys index points and multiplies the bill. Here max costs 16× low for +9.2 on the index — 2.3 points per doubling of spend. Compare that with other models.

low $0.345max $5.46

How its index has moved

29 Sept 2026 to 9 Oct 2026
546474
  • max
  • xhigh
  • high
  • medium

The index is recomputed from scratch every day, so a line moves when a new benchmark result lands, when a publisher revises a score, or when the models it is normalised against change. Early movement usually means the score is still settling.

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
GPQA Diamond––––95.6%#2–Epoch AI Benchmarking Hub
LMArena Hard Prompts–––1507#19––LMArena
LiveBench Reasoningnot in index–––86.8%#2991.6%#6–LiveBench
Humanity's Last Exam (AA)not in index36.2%#10539.8%#8245.8%#4450.0%#2355.0%#9–Artificial Analysis
CritPt11.4%#10316.9%#7724.6%#4431.1%#931.4%#7–Artificial Analysis
MysteryMechanismnot in indexeffort not stated–––––49.1%#3Vals AI
Mystery Game Puzzlesnot in index––––65.0%#4–Epoch AI Benchmarking Hub

Coding

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
SciCode49.1%#10652.9%#7753.7%#6557.3%#2661.0%#5–SciCode
FrontierCode–––52.1%#5––Cognition
LMArena Coding–––1532#18––LMArena
LMArena WebDev––1715#61774#3––LMArena
LiveBench Codingnot in index–––88.9%#491.4%#1–LiveBench
SciCode (AA)not in index49.1%#11352.9%#7753.7%#7057.3%#2761.0%#5–Artificial Analysis
IOIeffort not stated–––––83.1%#9Vals AI
Code Migrationnot in indexeffort not stated–––––69.8%#1Vals AI
Vibe Code Bench v1.1not in indexeffort not stated–––––92.4%#1Vals AI
CursorBenchnot in index35.8%#3739.2%#2947.8%#1053.1%#555.5%#4–Cursor
ALE-Benchnot in index––1819.1#10–––Epoch AI Benchmarking Hub
FrontierSWEnot in index––––61.9%#3–Epoch AI Benchmarking Hub
Terminal-Bench 4.0 (AA)not in index20.7%#5729.8%#4343.9%#2357.1%#663.6%#1–Artificial Analysis

Agents & tools

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
Terminal-Bench––––61.8%#14–Terminal-Bench
APEX-Agents–44.6%#35––75.5%#2–Mercor
LMArena Agent––––12#4–LMArena
LiveBench Agentic Codingnot in index–––39.3%#6256.3%#27–LiveBench
Terminal-Bench 2.1 (Vals)effort not stated–––––83.2%#6Vals AI
GDPval-AA v2.1not in index33.9%#11841.2%#9252.5%#3761.5%#667.0%#2–Artificial Analysis
Analyst Agent (AA)not in index––––57.5%#2–Artificial Analysis
CyberBenchnot in indexeffort not stated–––––59.6%#30Vals AI
Finance Agent v2not in indexeffort not stated–––––58.1%#10Vals AI
Harvey's Legal Agent Benchmarknot in indexeffort not stated–––––2.9%#36Vals AI
Legal Research Benchnot in indexeffort not stated–––––48.1%#8Vals AI
SREBenchnot in indexeffort not stated–––––30.1%#6Vals AI
Tax Agent Benchnot in indexeffort not stated–––––73.4%#4Vals AI
Terminal-Bench 4.0 (Vals)not in indexeffort not stated–––––64.1%#2Vals AI
Terminal-Bench Sciencenot in indexeffort not stated–––––38.6%#5Vals AI
AutomationBenchnot in index–––65.5%#1271.8%#2–Artificial Analysis
GDP.pdfnot in index–––24.6%#1625.8%#14–Artificial Analysis
Harvey LABnot in index––––2.8%#11–Harvey
AA-Briefcase v1.1not in index–––1752#41823#1–Artificial Analysis

Maths

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
FrontierMath Tiers 1–3––––88.8%#7–Epoch AI Benchmarking Hub
FrontierMath Tier 4––––80.5%#13–Epoch AI Benchmarking Hub
OTIS Mock AIME––––100.0%#1–Epoch AI Benchmarking Hub
ProofBench––––100.0%#1–Vals AI
LiveBench Mathematicsnot in index–––96.7%#696.1%#10–LiveBench
LMArena Maths–––1512#8––LMArena

Knowledge

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
SimpleQA Verified––––46.5%#34–Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index–––78.6%#2159.5%#61–LiveBench
AA-Omniscience19.4#6820.1#6520.9#6023.5#5232.3#26–Artificial Analysis
BioMysteryBenchnot in indexeffort not stated–––––81.1%#1Vals AI
Excel Modeling Benchmarknot in indexeffort not stated–––––75.7%#3Vals AI
MedCodenot in indexeffort not stated–––––52.9%#12Vals AI
MedScribenot in indexeffort not stated–––––91.1%#3Vals AI
Public Benefits Benchnot in indexeffort not stated–––––67.2%#12Vals AI
AA-Omniscience: accuracynot in index46.3%#8047.2%#7552.0%#5353.0%#4854.0%#42–Artificial Analysis
AA-Omniscience: non-hallucinationnot in index49.8%#11048.8%#11635.4%#18137.1%#17153.0%#100–Artificial Analysis

Instruction following

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
LiveBench Languagenot in index–––83.4%#1978.0%#38–LiveBench
LiveBench Instruction Followingnot in index–––70.5%#2756.8%#61–LiveBench
LMArena Instruction Followingnot in index–––1489#13––LMArena

Human preference

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
LMArena Text–––1476#31––LMArena
LMArena Creative Writingnot in index–––1463#24––LMArena
LMArena Multi-turnnot in index–––1484#31––LMArena
LMArena Longer Queriesnot in index–––1503#12––LMArena

Multimodal

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
LMArena Vision–––1289#28––LMArena
SAGEnot in indexeffort not stated–––––51.8%#11Vals AI

Long context

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
AA-LCR76.0%#14076.3%#13778.0%#11479.7%#8382.7%#33–Artificial Analysis
MLCRnot in index––––75.0%#1–Artificial Analysis

Composite

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
Epoch Capabilities Indexnot in indexeffort not stated–––––165.0#4Epoch AI Benchmarking Hub
LiveBench–––77.8%#2175.7%#35–LiveBench
AA Intelligence Index v4.3.235.9#8140.8#5646.8#2851.9#1056#2–Artificial Analysis
Vals Indexnot in indexeffort not stated–––––67.0#2Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Google Vertex121 tok/s2.26 s$2.20$11.001M–
Anthropic111 tok/s2.03 s$2.00$10.001M–
Azure110 tok/s1.19 s$2.00$10.001M–
Google Vertex100 tok/s0.57 s$2.20$11.001M–
Claude Platform on AWS98 tok/s2.64 s$2.00$10.001M–
Google Vertex95 tok/s1.80 s$2.00$10.001M–
Amazon Bedrock83 tok/s2.64 s$2.00$10.001M–

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.100 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0038$0.00328.2 min
Summarise a 30-page report12,000 / 600$0.030$0.0138.2 min
Code edit6,000 / 1,500$0.027$0.0188.3 min
Agentic coding session60,000 / 4,000$0.160$0.0748.6 min
Structured extraction2,000 / 200$0.0060$0.00328.1 min

See also

Data as of 11 Oct 2026. Compare these configurations.

Cite as: BenchLeader, “Claude Sonnet 5.5: benchmarks, pricing, speed and rank”, https://www.benchleader.com/models/claude-sonnet-5-5, data as of 11 Oct 2026.