BenchLeader

Deepseek v4 Pro vs GPT-5.5 Pro

Verdict
  • GPT-5.5 Pro (xhigh) leads on quality: 67.1 vs 60.8.
  • Deepseek v4 Pro (high) is stronger in coding, human preference.
  • GPT-5.5 Pro (xhigh) is stronger in maths, reasoning.
MetricDeepseek v4 Pro (high)GPT-5.5 Pro (xhigh)
BenchLeader Index60.867.1
Coding score59.3
Human preference score66.1
Maths score66.174.5
Reasoning score65.980.1
Blended price $/M$67.50
Output speed4 tok/s
Time to first answer3.5 s
Context window1.1M
GPQA Diamond90.9%
FrontierMath Tiers 1–387.7%
FrontierMath Tier 478.0%
OTIS Mock AIME95.6%
SciCode46.4%
WeirdML46.5%
LMArena Text1463
LMArena Hard Prompts1482
LMArena Coding1503
LMArena WebDev1581
ARC-AGI-195.0%
ARC-AGI-284.2%
CritPt30.6%
LMArena Maths1469
LMArena Creative Writing1442
LMArena Instruction Following1457
LMArena Multi-turn1469
LMArena Longer Queries1474
Chess Puzzles13.0%
CL-bench Life13.5%
Surface Evolver Bench40.0%
LMCA53.9%
DTBench96.0%
ALE-Bench1006.1

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

Deepseek v4 Pro vs GPT-5.5 Pro: questions

Is Deepseek v4 Pro better than GPT-5.5 Pro?
GPT-5.5 Pro (xhigh) leads on quality: 67.1 vs 60.8. The BenchLeader Index combines every independent quality benchmark; GPT-5.5 Pro (xhigh) is ahead overall as of 2026-09-19, but check the category scores for your use.