BenchLeader

Deepseek v4 Pro vs GPT-5.6 Terra

Verdict
  • GPT-5.6 Terra (xhigh) leads on quality: 62.6 vs 60.4.
  • Deepseek v4 Pro (high) is stronger in reasoning.
  • GPT-5.6 Terra (xhigh) is stronger in coding, human preference, maths, agents & tools, composite, instruction following, knowledge, long context, multimodal.
MetricDeepseek v4 Pro (high)GPT-5.6 Terra (xhigh)
BenchLeader Index60.462.6
Coding score59.361.5
Human preference score66.066.5
Maths score66.271.0
Reasoning score65.958.6
Agents & tools score66.9
Composite score70.6
Instruction following score62.1
Knowledge score59.9
Long context score63.5
Multimodal score62.4
Blended price $/M$4.50
Output speed97 tok/s
Time to first answer24.9 s
Context window1.1M
GPQA Diamond90.9%
OTIS Mock AIME95.6%
SimpleBench48.9%
SciCode46.4%51.6%
WeirdML46.5%
ProofBench74.0%
LMArena Text14621466
LMArena Hard Prompts14821489
LMArena Coding15051518
LMArena WebDev15811521
LMArena Vision1269
LMArena Agent1.5
AA Intelligence Index38.2
IFBench66.3%
AA-LCR79.0%
MMMU-Pro79.5%
AA-Omniscience-3.0
Terminal-Bench Hard62.9%
GPQA Diamond (AA)90.8%
Humanity's Last Exam (AA)41.9%
SciCode (AA)52.3%
τ²-Bench Telecom (AA)80.4%
LiveCodeBench85.9%
MMLU-Pro86.7%
CorpFin65.3%
TaxEval76.2%
GPQA Diamond (Vals)90.9%
ARC-AGI-194.0%
ARC-AGI-274.2%
ARC-AGI-30.7%

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

Deepseek v4 Pro vs GPT-5.6 Terra: questions

Is Deepseek v4 Pro better than GPT-5.6 Terra?
GPT-5.6 Terra (xhigh) leads on quality: 62.6 vs 60.4. The BenchLeader Index combines every independent quality benchmark; GPT-5.6 Terra (xhigh) is ahead overall as of 2026-09-13, but check the category scores for your use.
Is Deepseek v4 Pro better than GPT-5.6 Terra for coding?
GPT-5.6 Terra scores higher in coding (62 vs 59 on the category index, where 50 is average).