BenchLeader

DeepSeek V4 Pro vs GPT-5.4

Verdict
  • GPT-5.4 (xhigh) leads on quality: 65.4 vs 64.2.
  • DeepSeek V4 Pro (max) is stronger in agents & tools, composite, instruction following.
  • GPT-5.4 (xhigh) is stronger in coding, knowledge, long context, maths, reasoning, multimodal.
  • DeepSeek V4 Pro (max) is 10× cheaper ($0.544 vs $5.63 per 1M blended).
  • GPT-5.4 (xhigh) streams 1.5× faster (121 vs 81 tokens per second).
MetricDeepSeek V4 Pro (max)GPT-5.4 (xhigh)
BenchLeader Index64.265.4
Agents & tools score70.263.4
Coding score58.866.6
Composite score75.069.4
Instruction following score74.372.1
Knowledge score58.160.6
Long context score66.667.5
Maths score53.766.4
Reasoning score70.972.4
Multimodal score62.9
Blended price $/M$0.544$5.63
Output speed81 tok/s121 tok/s
Time to first answer55.6 s141.3 s
Context window1M1.1M
GPQA Diamond89.7%93.3%
FrontierMath Tiers 1–345.3%78.6%
FrontierMath Tier 42.4%49.0%
OTIS Mock AIME96.7%95.3%
SWE-bench Verified (Epoch)77.6%
SimpleQA Verified47.0%45.1%
Humanity's Last Exam36.2%
SciCode50.0%56.6%
WeirdML48.9%77.7%
APEX-Agents36.0%
ProofBench16.0%56.0%
GSO-Bench31.4%
LiveBench78.0%
LiveBench Reasoning88.1%
LiveBench Coding77.5%
LiveBench Agentic Coding53.8%
LiveBench Mathematics94.2%
LiveBench Data Analysis79.3%
LiveBench Language82.6%
LiveBench Instruction Following70.2%
AA Intelligence Index36.339.0
IFBench76.5%74.0%
AA-LCR80.3%82.0%
MMMU-Pro78.4%
AA-Omniscience0.85.8
Terminal-Bench Hard46.2%57.6%
GPQA Diamond (AA)92.8%92.0%
Humanity's Last Exam (AA)41.0%43.7%
SciCode (AA)51.0%
τ²-Bench Telecom (AA)96.2%87.1%
AIME (Vals)96.7%
LiveCodeBench87.5%84.1%
MMLU-Pro87.3%87.5%
LegalBench80.3%86.0%
CorpFin61.4%65.3%
TaxEval72.1%74.0%
MedQA96.1%
SWE-bench (Vals)77.4%78.2%
GPQA Diamond (Vals)89.4%91.7%
Vals Index42.9
AIME 202696.7%99.2%
HMMT February 202693.9%97.7%
MathArena Apex28.1%54.2%
SWE-Bench Pro59.1%
MCP Atlas70.6%
ARC-AGI-193.7%
ARC-AGI-274.0%
CritPt18.0%23.4%
GDPval (AA)49.7%40.3%
τ²-Bench Banking (AA)39.6%39.6%
ITBench SRE (AA)38.3%
Analyst Agent (AA)18.8%
APEX-Agents (AA)24.3%33.3%
CaseLaw v259.4%63.8%
Code Migration26.2%35.0%
Excel Modeling Benchmark51.6%
Finance Agent v244.1%
Harvey's Legal Agent Benchmark3.8%0.0%
Legal Research Bench23.1%
MedCode40.5%41.3%
MedScribe75.1%77.5%
MMMU-Pro (Vals)87.5%
MortgageTax68.3%
ProgramBench0.0%0.0%
Public Benefits Bench62.9%
SAGE43.3%
SkillsBench51.3%51.7%
Vibe Code Bench v1.149.9%67.4%
SWE-Bench Pro (private)43.4%
SWE Atlas: Codebase QnA40.8%
SWE Atlas: Refactoring44.3%
SWE Atlas: Test Writing44.4%
Chess Puzzles20.0%44.0%
EBR-bench25.4%
Mystery Game Puzzles37.0%
CL-bench27.9%
CL-bench Life21.7%
METR Time Horizons74.3%
DeepSWE51.8%
LMCA41.2%52.0%
DTBench90.7%94.4%

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

DeepSeek V4 Pro vs GPT-5.4: questions

Is DeepSeek V4 Pro better than GPT-5.4?
GPT-5.4 (xhigh) leads on quality: 65.4 vs 64.2. The BenchLeader Index combines every independent quality benchmark; GPT-5.4 (xhigh) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is DeepSeek V4 Pro better than GPT-5.4 for coding?
GPT-5.4 scores higher in coding (67 vs 59 on the category index, where 50 is average).
Is DeepSeek V4 Pro better than GPT-5.4 for agentic tasks?
DeepSeek V4 Pro scores higher in agentic tasks (70 vs 63 on the category index, where 50 is average).
Which is cheaper, DeepSeek V4 Pro or GPT-5.4?
DeepSeek V4 Pro is cheaper: $0.544 against $5.63 per million tokens, blended at three input tokens per output token.
Which is faster, DeepSeek V4 Pro or GPT-5.4?
GPT-5.4 streams faster: 121 against 81 output tokens per second.
Which has the larger context window?
GPT-5.4 accepts more context: 1.1M against 1M tokens.