BenchLeader

Claude Opus 5.5 vs Deepseek v4 Pro High

Verdict
  • Claude Opus 5.5 (thinking) leads on quality: 70.7 vs 60.8.
  • Claude Opus 5.5 (thinking) is stronger in composite, knowledge, long context, multimodal, reasoning.
  • Deepseek v4 Pro High (high) is stronger in coding, human preference, maths.
MetricClaude Opus 5.5 (thinking)Deepseek v4 Pro High (high)
BenchLeader Index70.760.8
Composite score95.0
Knowledge score85.0
Long context score68.5
Multimodal score71.9
Reasoning score95.066.0
Coding score59.0
Human preference score66.2
Maths score66.2
Blended price $/M$8.00
Output speed48 tok/s
Time to first answer4.2 s
Context window1M
GPQA Diamond90.9%
OTIS Mock AIME95.6%
SciCode46.4%
WeirdML46.5%
LMArena Text1463
LMArena Hard Prompts1482
LMArena Coding1503
LMArena WebDev1581
AA Intelligence Index57.6
AA-LCR84.7%
MMMU-Pro87.7%
AA-Omniscience46.4
Humanity's Last Exam (AA)61.4%
SciCode (AA)66.9%
CritPt31.7%
GDPval (AA)67.3%
LMArena Maths1469
LMArena Creative Writing1442
LMArena Instruction Following1457
LMArena Multi-turn1469
LMArena Longer Queries1474
Chess Puzzles13.0%
CL-bench Life13.5%
Surface Evolver Bench40.0%
ALE-Bench1006.1
Terminal-Bench 4.0 (AA)59.6%
AutomationBench69.5%
GDP.pdf26.2%
Harvey LAB91.2%
AA-Omniscience: accuracy66.2%
AA-Omniscience: non-hallucination41.4%
AA-Briefcase1822

Data as of 2026-09-22. Best configuration of each model; every score links to its source on the model pages.

Claude Opus 5.5 vs Deepseek v4 Pro High: questions

Is Claude Opus 5.5 better than Deepseek v4 Pro High?
Claude Opus 5.5 (thinking) leads on quality: 70.7 vs 60.8. The BenchLeader Index combines every independent quality benchmark; Claude Opus 5.5 (thinking) is ahead overall as of 2026-09-22, but check the category scores for your use.