BenchLeader

Deepseek v4 Pro vs GPT-5.2

Verdict
  • Deepseek v4 Pro (high) and GPT-5.2 are level on quality (60.4 vs 60.3).
  • Deepseek v4 Pro (high) is stronger in coding, maths, reasoning.
  • GPT-5.2 is stronger in human preference, agents & tools, composite, instruction following, knowledge, long context, multimodal.
MetricDeepseek v4 Pro (high)GPT-5.2
BenchLeader Index60.460.3
Coding score59.353.7
Human preference score66.067.8
Maths score66.2
Reasoning score65.961.3
Agents & tools score60.4
Composite score61.6
Instruction following score65.4
Knowledge score58.8
Long context score65.6
Multimodal score59.9
Blended price $/M$4.81
Output speed72 tok/s
Time to first answer90.8 s
Context window400k
GPQA Diamond90.9%
OTIS Mock AIME95.6%
Humanity's Last Exam27.8%
Terminal-Bench64.9%
SimpleBench45.8%
SciCode46.4%
Remote Labor Index2.1%
WeirdML46.5%
APEX-Agents23.0%
Epoch Capabilities Index153.5
LMArena Text14621476
LMArena Hard Prompts14821497
LMArena Coding15051515
LMArena WebDev15811417
LMArena Vision1268
AA Intelligence Index30.4
IFBench75.4%
AA-LCR82.7%
AA-Omniscience-0.9
Terminal-Bench Hard47.0%
GPQA Diamond (AA)90.3%
Humanity's Last Exam (AA)37.7%
τ²-Bench Telecom (AA)84.8%
SWE-Bench Pro29.9%
VISTA46.6%
MultiNRC42.2%
TutorBench53.5%
Kagi LLM Benchmark73.3%
SWE-bench Verified (bash only)69.0%
ARC-AGI-194.5%
ARC-AGI-272.9%
BFCL Overall55.9%

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

Deepseek v4 Pro vs GPT-5.2: questions

Is Deepseek v4 Pro better than GPT-5.2?
Deepseek v4 Pro (high) and GPT-5.2 are level on quality (60.4 vs 60.3). The BenchLeader Index combines every independent quality benchmark; Deepseek v4 Pro (high) is ahead overall as of 2026-09-13, but check the category scores for your use.
Is Deepseek v4 Pro better than GPT-5.2 for coding?
Deepseek v4 Pro scores higher in coding (59 vs 54 on the category index, where 50 is average).