BenchLeader

Deepseek v4 Pro vs GPT-6 Astra

Verdict
  • GPT-6 Astra (max) leads on quality: 70.9 vs 60.4.
  • Deepseek v4 Pro (high) is stronger in human preference.
  • GPT-6 Astra (max) is stronger in coding, maths, reasoning, agents & tools, composite, knowledge.
MetricDeepseek v4 Pro (high)GPT-6 Astra (max)
BenchLeader Index60.470.9
Coding score59.478.6
Human preference score65.8
Maths score66.277.2
Reasoning score66.073.6
Agents & tools score68.2
Composite score72.9
Knowledge score80.0
Blended price $/M$20.00
Output speed27 tok/s
Time to first answer337.3 s
Context window1.1M
GPQA Diamond90.9%95.8%
FrontierMath Tiers 1–393.7%
FrontierMath Tier 497.6%
OTIS Mock AIME95.6%100.0%
SimpleQA Verified75.6%
Terminal-Bench58.2%
SciCode46.4%56.5%
WeirdML46.5%
FrontierCode53.3%
LMArena Text1460
LMArena Hard Prompts1482
LMArena Coding1505
LMArena WebDev15801796
LMArena Agent12.5
LiveBench82.2%
LiveBench Reasoning92.7%
LiveBench Coding80.4%
LiveBench Agentic Coding57.3%
LiveBench Mathematics96.8%
LiveBench Data Analysis83.0%
LiveBench Language89.4%
Terminal-Bench 2.1 (Vals)87.3%
Vals Index66.6
ARC-AGI-197.5%
ARC-AGI-295.0%
ARC-AGI-398.6%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Deepseek v4 Pro vs GPT-6 Astra: questions

Is Deepseek v4 Pro better than GPT-6 Astra?
GPT-6 Astra (max) leads on quality: 70.9 vs 60.4. The BenchLeader Index combines every independent quality benchmark; GPT-6 Astra (max) is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Deepseek v4 Pro better than GPT-6 Astra for coding?
GPT-6 Astra scores higher in coding (79 vs 59 on the category index, where 50 is average).