BenchLeader

Deepseek v4 Pro vs GPT-5.3-Codex

Verdict
  • Deepseek v4 Pro (high) and GPT-5.3-Codex are level on quality (60.4 vs 60.6).
  • Deepseek v4 Pro (high) is stronger in coding, human preference, maths, reasoning.
  • GPT-5.3-Codex is stronger in agents & tools, composite, instruction following, knowledge, long context, multimodal.
MetricDeepseek v4 Pro (high)GPT-5.3-Codex
BenchLeader Index60.460.6
Coding score59.354.7
Human preference score66.0
Maths score66.2
Reasoning score65.9
Agents & tools score61.7
Composite score64.0
Instruction following score70.1
Knowledge score65.9
Long context score66.0
Multimodal score61.6
Blended price $/M$4.81
Output speed124 tok/s
Time to first answer48.3 s
Context window400k
GPQA Diamond90.9%
OTIS Mock AIME95.6%
Terminal-Bench78.4%
SciCode46.4%
WeirdML46.5%79.3%
APEX-Agents31.8%
Epoch Capabilities Index156.6
LMArena Text1462
LMArena Hard Prompts1482
LMArena Coding1505
LMArena WebDev15811409
AA Intelligence Index32.5
IFBench75.4%
AA-LCR83.3%
MMMU-Pro78.5%
AA-Omniscience10.9
Terminal-Bench Hard53.0%
GPQA Diamond (AA)91.5%
Humanity's Last Exam (AA)42.5%
τ²-Bench Telecom (AA)86.0%
HiL-Bench4.3%

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

Deepseek v4 Pro vs GPT-5.3-Codex: questions

Is Deepseek v4 Pro better than GPT-5.3-Codex?
Deepseek v4 Pro (high) and GPT-5.3-Codex are level on quality (60.4 vs 60.6). The BenchLeader Index combines every independent quality benchmark; GPT-5.3-Codex is ahead overall as of 2026-09-13, but check the category scores for your use.
Is Deepseek v4 Pro better than GPT-5.3-Codex for coding?
Deepseek v4 Pro scores higher in coding (59 vs 55 on the category index, where 50 is average).