BenchLeader
AnthropicReasoning modelNewReleased 7 Oct 2026Compare with another model

Claude Haiku 5.5

Reasoning effort

Claude Haiku 5.5 is an Anthropic proprietary reasoning model, released 7 Oct 2026. Its best configuration (high reasoning effort) ranks #92 of 760 on the BenchLeader Index at 61.3 ±3.2, in the upper half. The ± is the point: 176 other configurations score within that range, so they and this one cannot be told apart on quality alone — price and speed are what separate them. It scores highest in composite (72) and lowest in coding (63). At $0.200 per million tokens blended it is cheaper than most ranked models. Output speed of 174 tokens per second puts it in the fastest quarter, with a first answer in 22.6 s. It has been measured at 6 reasoning-effort settings; this summary describes the best-scoring one, and the tabs above switch between them. Last measured 8 Oct 2026.

Blended price
$0.200/M
$0.100 in · $0.500 out
Output speed
174 tok/s
measured by Artificial Analysis
First answer
23 s
first token 1.75 s
Context
1M
Full answer
26 s
median, reasoning included
Cost per run
$0.079
one full Intelligence Index run
Released
7 Oct 2026
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index61
  2. Reasoning72
  3. Coding63
  4. Maths66
  5. Knowledge64
  6. Long context64
  7. Composite72

Versions

Anthropic has shipped 6 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
NewClaude Haiku 5.5highthis page7 Oct 202661.3#92
Claude Haiku 4.5no reasoning15 Oct 202548.5#409
Claude 3.5 Haiku4 Nov 202439.8#650
Claude 3 Haiku7 Mar 202437.1#714
Claude Haiku 3–––
Claude Haiku 4thinking–––

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Not stated” collects results from publishers that did not say which setting they used; for a reasoning model that is usually its thinking mode, but we do not assume it. Pick a setting here or at the top of the page to see its price, speed and category scores.

EffortIndexRankSpeedFirst answerChat reply costCategories
low––178 tok/s9.72 s$0.0002Composite 63 · Knowledge 63 · Long context 61 · Reasoning 56
medium––166 tok/s13 s$0.0002Composite 69 · Knowledge 64 · Long context 64 · Reasoning 63
highbest61.3#92174 tok/s23 s$0.0002Coding 63 · Composite 72 · Knowledge 64 · Long context 64 · Maths 66 · Reasoning 72
xhigh61.1#94188 tok/s57 s$0.0002Composite 59 · Knowledge 64 · Long context 64 · Maths 67 · Reasoning 79
max57.9#157240 tok/s289 s$0.0002Composite 54 · Knowledge 51 · Long context 67 · Maths 62 · Reasoning 69
not stated––146 tok/s1.75 s$0.0002Instruction following 64

What more thinking costs

Turning the effort up buys index points and multiplies the bill. Here max costs 2.7× high for -3.4 on the index — -2.4 points per doubling of spend. Compare that with other models.

high $0.079max $0.213

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
GPQA Diamond––––89.7%#50–Epoch AI Benchmarking Hub
LiveBench Reasoningnot in index–––81.2%#4984.6%#39–LiveBench
Humanity's Last Exam (AA)not in index27.0%#16633.8%#12137.3%#9942.7%#5844.4%#48–Artificial Analysis
CritPt9.1%#11812.9%#9718.6%#6322.6%#5118.9%#62–Artificial Analysis
Mystery Game Puzzlesnot in index––––30.0%#31–Epoch AI Benchmarking Hub

Coding

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
LMArena WebDev––1587#29–––LMArena
LiveBench Codingnot in index–––76.4%#4577.9%#34–LiveBench
SciCode (AA)not in index49.2%#11149.0%#11448.7%#11551.7%#8555.0%#50–Artificial Analysis
Terminal-Bench 4.0 (AA)not in index12.6%#7315.2%#6521.7%#5429.3%#4432.8%#38–Artificial Analysis

Agents & tools

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
LiveBench Agentic Codingnot in index–––51.4%#3957.2%#21–LiveBench
GDPval-AA v2.1not in index31.1%#12938.7%#9945.9%#6850.5%#4355.9%#24–Artificial Analysis
AutomationBenchnot in index22.9%#4528.6%#4433.7%#4336.0%#4035.4%#42–Artificial Analysis
GDP.pdfnot in index11.2%#4315.2%#3917.2%#3418.2%#3220.8%#25–Artificial Analysis
Harvey LABnot in index1.7%#130.6%#171.1%#152.8%#110.6%#17–Harvey
AA-Briefcase v1.1not in index1112#431372#381442#321533#211577#13–Artificial Analysis

Maths

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
FrontierMath Tiers 1–3––––75.1%#18–Epoch AI Benchmarking Hub
FrontierMath Tier 4––––46.3%#25–Epoch AI Benchmarking Hub
OTIS Mock AIME––97.2%#3298.9%#16––Epoch AI Benchmarking Hub
LiveBench Mathematicsnot in index–––92.8%#2686.4%#49–LiveBench

Knowledge

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
SimpleQA Verified––––23.8%#67–Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index–––72.8%#4351.8%#66–LiveBench
AA-Omniscience3.3#1104.5#1055.8#996.0#9710.7#89–Artificial Analysis
AA-Omniscience: accuracynot in index33.3%#15634.0%#15234.8%#14935.0%#14536.4%#142–Artificial Analysis
AA-Omniscience: non-hallucinationnot in index55.1%#9355.2%#9155.5%#8955.6%#8859.6%#74–Artificial Analysis

Instruction following

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
LiveBench Languagenot in index–––63.5%#6262.6%#64–LiveBench
LiveBench Instruction Followingnot in index–––66.5%#3954.7%#64–LiveBench
MultiChallengeeffort not stated–––––67.7%#6Scale AI SEAL

Long context

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
AA-LCR72.0%#18377.3%#12477.3%#12478.3%#10682.7%#33–Artificial Analysis
MLCRnot in index––––63.3%#5–Artificial Analysis

Composite

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
LiveBench–––72.1%#4967.9%#58–LiveBench
AA Intelligence Index v4.3.229.4#12034.5#8637.8#7641.3#5343.4#43–Artificial Analysis

Safety & honesty

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
FORTRESSnot in indexeffort not stated–––––9.5%#66Scale AI SEAL

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Google Vertex218 tok/s0.72 s$0.110$0.5501M–
Google Vertex173 tok/s1.48 s$0.110$0.5501M–
Anthropic167 tok/s1.69 s$0.100$0.5001M–
Google Vertex156 tok/s1.43 s$0.100$0.5001M–
Amazon Bedrock149 tok/s2.70 s$0.100$0.5001M–
Claude Platform on AWS143 tok/s1.94 s$0.100$0.5001M–
Azure124 tok/s0.74 s$0.100$0.5001M–
Amazon Bedrock74 tok/s0.82 s$0.110$0.5501M–
Amazon Bedrock–0.60 s$0.110$0.5501M–

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.010 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0002$0.000224.4 s
Summarise a 30-page report12,000 / 600$0.0015$0.000726.1 s
Code edit6,000 / 1,500$0.0014$0.000931.3 s
Agentic coding session60,000 / 4,000$0.0080$0.004045.6 s
Structured extraction2,000 / 200$0.0003$0.000223.8 s

See also

Data as of 11 Oct 2026. Compare these configurations.

Cite as: BenchLeader, “Claude Haiku 5.5: benchmarks, pricing, speed and rank”, https://www.benchleader.com/models/claude-haiku-5-5, data as of 11 Oct 2026.