BenchLeader

Claude Opus 5 vs Claude Opus 5.5

Verdict
  • Claude Opus 5 (high) and Claude Opus 5.5 (thinking) are level on quality (69.8 vs 70.7).
  • Claude Opus 5 (high) is stronger in agents & tools, coding, human preference, maths.
  • Claude Opus 5.5 (thinking) is stronger in composite, knowledge, long context, multimodal, reasoning.
  • Claude Opus 5.5 (thinking) is 1.3× cheaper ($8.00 vs $10.00 per 1M blended).
MetricClaude Opus 5 (high)Claude Opus 5.5 (thinking)
BenchLeader Index69.870.7
Agents & tools score66.0
Coding score71.4
Composite score87.995.0
Human preference score69.8
Knowledge score79.185.0
Long context score65.668.5
Maths score72.3
Multimodal score67.371.9
Reasoning score75.295.0
Blended price $/M$10.00$8.00
Output speed55 tok/s48 tok/s
Time to first answer16.1 s4.2 s
Context window1M1M
Terminal-Bench50.3%
OSWorld-Verified 2.029.0%
SciCode54.3%
WeirdML91.6%
LMArena Text1493
LMArena Hard Prompts1517
LMArena Coding1533
LMArena WebDev1660
LMArena Vision1321
LMArena Agent10.2
AA Intelligence Index48.157.6
AA-LCR79.0%84.7%
MMMU-Pro82.4%87.7%
AA-Omniscience33.746.4
GPQA Diamond (AA)93.7%
Humanity's Last Exam (AA)52.8%61.4%
SciCode (AA)55.4%66.9%
ARC-AGI-197.5%
ARC-AGI-288.3%
ARC-AGI-330.2%
CritPt28.3%31.7%
GDPval (AA)54.0%67.3%
τ²-Bench Banking (AA)44.7%
LMArena Maths1525
LMArena Creative Writing1474
LMArena Instruction Following1499
LMArena Multi-turn1484
LMArena Longer Queries1508
LMArena Document1490
DeepSWE72.8%
LMCA63.7%
DTBench97.9%
CursorBench44.7%
ALE-Bench2164.6
Terminal-Bench 4.0 (AA)46.0%59.6%
Terminal-Bench 2.1 (AA)87.6%
AutomationBench53.6%69.5%
GDP.pdf19.6%26.2%
MLCR59.4%
Harvey LAB91.2%
AA-Omniscience: accuracy58.9%66.2%
AA-Omniscience: non-hallucination38.8%41.4%
AA-Briefcase15731822

Data as of 2026-09-22. Best configuration of each model; every score links to its source on the model pages.

Claude Opus 5 vs Claude Opus 5.5: questions

Is Claude Opus 5 better than Claude Opus 5.5?
Claude Opus 5 (high) and Claude Opus 5.5 (thinking) are level on quality (69.8 vs 70.7). The BenchLeader Index combines every independent quality benchmark; Claude Opus 5.5 (thinking) is ahead overall as of 2026-09-22, but check the category scores for your use.
Which is cheaper, Claude Opus 5 or Claude Opus 5.5?
Claude Opus 5.5 is cheaper: $8.00 against $10.00 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Opus 5 or Claude Opus 5.5?
Claude Opus 5 streams faster: 55 against 48 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.