BenchLeader

Claude Opus 4.5 vs Deepseek v4 Pro

Verdict
  • Deepseek v4 Pro (high) leads on quality: 60.4 vs 59.4.
  • Claude Opus 4.5 (thinking) is stronger in agents & tools, coding, composite, instruction following, knowledge, long context, multimodal.
  • Deepseek v4 Pro (high) is stronger in maths, reasoning, human preference.
MetricClaude Opus 4.5 (thinking)Deepseek v4 Pro (high)
BenchLeader Index59.460.4
Agents & tools score69.7
Coding score60.559.3
Composite score60.0
Instruction following score53.5
Knowledge score60.3
Long context score62.5
Maths score64.066.2
Multimodal score57.0
Reasoning score58.665.9
Human preference score66.0
Blended price $/M$10.00
Output speed44 tok/s
Time to first answer13.7 s
Context window200k
GPQA Diamond90.9%
OTIS Mock AIME95.6%
SciCode46.4%
WeirdML46.5%
LMArena Text1462
LMArena Hard Prompts1482
LMArena Coding1505
LMArena WebDev1581
AA Intelligence Index29.1
IFBench58.0%
AA-LCR77.3%
MMMU-Pro74.0%
AA-Omniscience14
Terminal-Bench Hard47.0%
GPQA Diamond (AA)86.6%
Humanity's Last Exam (AA)30.1%
τ²-Bench Telecom (AA)89.5%
AIME (Vals)95.4%
LiveCodeBench83.7%
MMLU-Pro87.3%
LegalBench84.6%
CorpFin65.1%
TaxEval74.9%
MedQA95.9%
MGSM95.2%
SWE-bench (Vals)76.4%
GPQA Diamond (Vals)85.9%
MultiChallenge59.0%
PRBench Finance46.2%
PRBench Legal44.2%
VISTA46.4%
MultiNRC48.6%
TutorBench51.2%
Kagi LLM Benchmark80.2%
ARC-AGI-180.0%
ARC-AGI-237.6%

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

Claude Opus 4.5 vs Deepseek v4 Pro: questions

Is Claude Opus 4.5 better than Deepseek v4 Pro?
Deepseek v4 Pro (high) leads on quality: 60.4 vs 59.4. The BenchLeader Index combines every independent quality benchmark; Deepseek v4 Pro (high) is ahead overall as of 2026-09-13, but check the category scores for your use.
Is Claude Opus 4.5 better than Deepseek v4 Pro for coding?
Claude Opus 4.5 scores higher in coding (61 vs 59 on the category index, where 50 is average).