BenchLeader

Claude Opus 4.5 vs Claude Opus 5.5

Verdict
  • Claude Opus 5.5 (thinking) leads on quality: 70.7 vs 60.6.
  • Claude Opus 4.5 (thinking) is stronger in agents & tools, coding, instruction following, maths.
  • Claude Opus 5.5 (thinking) is stronger in composite, knowledge, long context, multimodal, reasoning.
  • Claude Opus 5.5 (thinking) is 1.3× cheaper ($8.00 vs $10.00 per 1M blended).
MetricClaude Opus 4.5 (thinking)Claude Opus 5.5 (thinking)
BenchLeader Index60.670.7
Agents & tools score74.5
Coding score60.6
Composite score64.795.0
Instruction following score54.8
Knowledge score60.685.0
Long context score64.768.5
Maths score64.0
Multimodal score57.871.9
Reasoning score55.795.0
Blended price $/M$10.00$8.00
Output speed49 tok/s48 tok/s
Time to first answer16.5 s4.2 s
Context window200k1M
AA Intelligence Index29.157.6
IFBench58.0%
AA-LCR77.3%84.7%
MMMU-Pro74.0%87.7%
AA-Omniscience1446.4
Terminal-Bench Hard47.0%
GPQA Diamond (AA)86.6%
Humanity's Last Exam (AA)30.1%61.4%
SciCode (AA)66.9%
τ²-Bench Telecom (AA)89.5%
AIME (Vals)95.4%
LiveCodeBench83.7%
MMLU-Pro87.3%
LegalBench84.6%
CorpFin65.1%
TaxEval74.9%
MedQA95.9%
MGSM95.2%
SWE-bench (Vals)76.4%
GPQA Diamond (Vals)85.9%
MultiChallenge59.0%
PRBench Finance46.2%
PRBench Legal44.2%
VISTA46.4%
MultiNRC48.6%
TutorBench51.2%
Kagi LLM Benchmark80.2%
ARC-AGI-180.0%
ARC-AGI-237.6%
CritPt4.6%31.7%
GDPval (AA)67.3%
CaseLaw v262.6%
MedCode49.2%
MedScribe85.3%
MMMU-Pro (Vals)83.0%
MortgageTax67.7%
SAGE52.1%
Terminal-Bench 2.0 (Vals)53.9%
Vibe Code Bench v1.120.6%
FORTRESS9.6%
MASK92.5%
Terminal-Bench 4.0 (AA)59.6%
AutomationBench69.5%
GDP.pdf26.2%
Harvey LAB91.2%
AA-Omniscience: accuracy46.6%66.2%
AA-Omniscience: non-hallucination39.0%41.4%
AA-Briefcase1822
DeepSearchQA24.0%

Data as of 2026-09-22. Best configuration of each model; every score links to its source on the model pages.

Claude Opus 4.5 vs Claude Opus 5.5: questions

Is Claude Opus 4.5 better than Claude Opus 5.5?
Claude Opus 5.5 (thinking) leads on quality: 70.7 vs 60.6. The BenchLeader Index combines every independent quality benchmark; Claude Opus 5.5 (thinking) is ahead overall as of 2026-09-22, but check the category scores for your use.
Which is cheaper, Claude Opus 4.5 or Claude Opus 5.5?
Claude Opus 5.5 is cheaper: $8.00 against $10.00 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Opus 4.5 or Claude Opus 5.5?
Claude Opus 4.5 streams faster: 49 against 48 output tokens per second.
Which has the larger context window?
Claude Opus 5.5 accepts more context: 1M against 200k tokens.