BenchLeader
AnthropicReasoning modelNewReleased 22 Sept 2026

Claude Opus 5.5

Reasoning effort

Claude Opus 5.5 is an Anthropic proprietary reasoning model, released 22 Sept 2026. Its best configuration (thinking reasoning effort) ranks #6 of 430 on the BenchLeader Index at 70.7 ±9.8, in the top ten. It scores highest in composite (95) and lowest in long context (69). At $8.00 per million tokens blended it is among the most expensive ranked models. Output speed of 48 tokens per second puts it slower than most, with a first token in 4.2 s. It has been measured at 5 reasoning-effort settings; this summary describes the best-scoring one, and the tabs above switch between them. Last measured 22 Sept 2026.

Blended price
$8.00/M
$4.00 in · $20.00 out
Output speed
48 tok/s
OpenRouter traffic, 7-day median; not yet measured by Artificial Analysis
First answer
4.25 s
Context
1M
Full answer
Cost per run
Released
22 Sept 2026
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index71
  2. Reasoning95
  3. Knowledge85
  4. Multimodal72
  5. Long context69
  6. Composite95

Versions

Anthropic has shipped 9 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
NewClaude Opus 5.5thinkingthis page22 Sept 202670.7#6
Claude Opus 5high24 Jul 202669.8#11
Claude Opus 4.8max28 May 202663.5#65
Claude Opus 4.716 Apr 202664.4#48
Claude Opus 4.65 Feb 202663.7#60
Claude Opus 4.5thinking24 Nov 202560.6#97
Claude Opus 4.1thinking5 Aug 202557.0#167
Claude Opus 4thinking22 May 202555.7#192
Claude 3 Opus4 Mar 202439.7#628

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Not stated” collects results from publishers that did not say which setting they used; for a reasoning model that is usually its thinking mode, but we do not assume it. Pick a setting here or at the top of the page to see its price, speed and category scores.

EffortIndexRankSpeedFirst answerChat reply costCategories
low65.2#3948 tok/s4.25 s$0.0076Composite 81 · Knowledge 81 · Long context 67 · Multimodal 69 · Reasoning 74
medium69.3#1548 tok/s4.25 s$0.0076Composite 92 · Knowledge 82 · Long context 68 · Multimodal 70 · Reasoning 92
high69.9#1048 tok/s4.25 s$0.0076Composite 95 · Knowledge 82 · Long context 68 · Multimodal 70 · Reasoning 95
xhigh70.3#948 tok/s4.25 s$0.0076Composite 95 · Knowledge 83 · Long context 69 · Multimodal 71 · Reasoning 95
thinkingbest70.7#648 tok/s4.25 s$0.0076Composite 95 · Knowledge 85 · Long context 69 · Multimodal 72 · Reasoning 95

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarklowmediumhighxhighthinkingSource
Humanity's Last Exam (AA)not in index48.3%#2454.7%#955.6%#657.5%#461.4%#1Artificial Analysis
CritPt17.7%#5127.7%#2030.9%#731.7%#231.7%#2Artificial Analysis

Coding

BenchmarklowmediumhighxhighthinkingSource
SciCode (AA)not in index58.6%#1859.3%#1160.4%#765.0%#266.9%#1Artificial Analysis
Terminal-Bench 4.0 (AA)not in index31.3%#2852.5%#856.6%#559.6%#159.6%#1Artificial Analysis

Agents & tools

BenchmarklowmediumhighxhighthinkingSource
GDPval (AA)not in index36.2%#8053.8%#2659.6%#866.0%#267.3%#1Artificial Analysis
AutomationBenchnot in index63.2%#1265.0%#969.5%#1Artificial Analysis
GDP.pdfnot in index28.8%#626.6%#1026.2%#12Artificial Analysis
Harvey LABnot in index90.9%#1291.2%#1091.2%#11Harvey
AA-Briefcasenot in index1705#31780#21822#1Artificial Analysis

Knowledge

BenchmarklowmediumhighxhighthinkingSource
AA-Omniscience38.9#1440.3#1340.6#1142.6#746.4#1Artificial Analysis
AA-Omniscience: accuracynot in index63.5%#964.5%#864.6%#765.4%#466.2%#2Artificial Analysis
AA-Omniscience: non-hallucinationnot in index32.4%#17431.6%#17832.4%#17334.3%#15941.4%#125Artificial Analysis

Multimodal

BenchmarklowmediumhighxhighthinkingSource
MMMU-Pro84.7%#1485.7%#785.8%#686.6%#387.7%#1Artificial Analysis

Long context

BenchmarklowmediumhighxhighthinkingSource
AA-LCR80.7%#4984.3%#882.7%#2684.7%#584.7%#5Artificial Analysis

Composite

BenchmarklowmediumhighxhighthinkingSource
AA Intelligence Index42.3#3551.2#853.6#356.0#257.6#1Artificial Analysis

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Anthropic67 tok/s4.69 s$4.00$20.001M
Amazon Bedrock49 tok/s4.35 s$4.00$20.001M
Claude Platform on AWS48 tok/s4.25 s$4.00$20.001M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.200 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0076$0.006510.5 s
Summarise a 30-page report12,000 / 600$0.060$0.02616.7 s
Code edit6,000 / 1,500$0.054$0.03735.5 s
Agentic coding session60,000 / 4,000$0.320$0.1491.5 min
Structured extraction2,000 / 200$0.012$0.00638.4 s

See also

Data as of 22 Sept 2026. Compare these configurations.

Cite as: BenchLeader, “Claude Opus 5.5: benchmarks, pricing, speed and rank”, https://www.benchleader.com/models/claude-opus-5-5, data as of 22 Sept 2026.