BenchLeader

Deepseek v4 Pro vs o3

Verdict
  • Deepseek v4 Pro (high) and o3 are level on quality (60.4 vs 60.5).
  • Deepseek v4 Pro (high) is stronger in human preference, reasoning.
  • o3 is stronger in coding, maths, agents & tools, composite, instruction following, knowledge, long context, multimodal.
MetricDeepseek v4 Pro (high)o3
BenchLeader Index60.460.5
Coding score59.362.0
Human preference score66.062.3
Maths score66.278.6
Reasoning score65.961.5
Agents & tools score68.7
Composite score49.6
Instruction following score63.8
Knowledge score56.2
Long context score60.9
Multimodal score56.9
Blended price $/M$3.50
Output speed130 tok/s
Time to first answer4.4 s
Context window200k
GPQA Diamond90.9%
OTIS Mock AIME95.6%
SciCode46.4%
WeirdML46.5%
Epoch Capabilities Index146.9
LMArena Text14621432
LMArena Hard Prompts14821441
LMArena Coding15051460
LMArena WebDev1581
LMArena Vision1214
AA Intelligence Index20.2
IFBench71.4%
AA-LCR74.7%
MMMU-Pro70.1%
AA-Omniscience-15.6
Terminal-Bench Hard37.1%
GPQA Diamond (AA)82.7%
Humanity's Last Exam (AA)20.1%
τ²-Bench Telecom (AA)80.7%
PRBench Finance47.7%
PRBench Legal48.6%
MMMU (validation)82.9%
MMMU-Pro (official)76.4%
Kagi LLM Benchmark67.6%
IFEval (HELM)86.9%
Omni-MATH (HELM)71.4%
WildBench (HELM)86.1%
MMLU-Pro (HELM)85.9%
GPQA Diamond (HELM)75.3%
HELM Capabilities mean81.1%
Aider Polyglot81.3%
SWE-bench Verified (bash only)58.4%
SWE-bench Verified (any scaffold)58.4%
BFCL Overall63.0%

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

Deepseek v4 Pro vs o3: questions

Is Deepseek v4 Pro better than o3?
Deepseek v4 Pro (high) and o3 are level on quality (60.4 vs 60.5). The BenchLeader Index combines every independent quality benchmark; o3 is ahead overall as of 2026-09-13, but check the category scores for your use.
Is Deepseek v4 Pro better than o3 for coding?
o3 scores higher in coding (62 vs 59 on the category index, where 50 is average).