BenchLeader

Deepseek v4 Pro vs Gemini 3.1 Pro

Verdict
  • Gemini 3.1 Pro leads on quality: 62.8 vs 60.4.
  • Deepseek v4 Pro (high) is stronger in human preference, maths.
  • Gemini 3.1 Pro is stronger in coding, reasoning, agents & tools, composite, instruction following, knowledge, long context, multimodal.
MetricDeepseek v4 Pro (high)Gemini 3.1 Pro
BenchLeader Index60.462.8
Coding score59.359.8
Human preference score66.059.8
Maths score66.260.7
Reasoning score65.969.8
Agents & tools score62.6
Composite score61.5
Instruction following score68.4
Knowledge score64.9
Long context score65.2
Multimodal score65.8
Blended price $/M$4.50
Output speed115 tok/s
Time to first answer23.4 s
Context window1.0M
GPQA Diamond90.9%94.1%
FrontierMath Tiers 1–359.6%
FrontierMath Tier 426.8%
OTIS Mock AIME95.6%95.6%
Humanity's Last Exam46.4%
Terminal-Bench80.2%
SimpleBench79.6%
SciCode46.4%58.9%
WeirdML46.5%72.1%
APEX-Agents33.5%
ProofBench26.0%
GSO-Bench22.6%
Epoch Capabilities Index155
LMArena Text14621487
LMArena Hard Prompts14821508
LMArena Coding15051520
LMArena WebDev15811447
LMArena Vision1295
LMArena Agent-5.6
AA Intelligence Index30.4
IFBench77.1%
AA-LCR82.0%
MMMU-Pro82.4%
AA-Omniscience31.9
Terminal-Bench Hard53.8%
GPQA Diamond (AA)94.1%
Humanity's Last Exam (AA)47.0%
SciCode (AA)58.7%
τ²-Bench Telecom (AA)95.6%
AIME 202698.3%
HMMT February 202694.7%
MathArena Apex60.9%
MultiChallenge71.4%
PRBench Finance41.9%
PRBench Legal44.0%
MultiNRC64.7%
HiL-Bench35.3%
TutorBench53.0%
EQ-Bench 41142
ARC-AGI-198.0%
ARC-AGI-277.1%
ARC-AGI-30.4%

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

Deepseek v4 Pro vs Gemini 3.1 Pro: questions

Is Deepseek v4 Pro better than Gemini 3.1 Pro?
Gemini 3.1 Pro leads on quality: 62.8 vs 60.4. The BenchLeader Index combines every independent quality benchmark; Gemini 3.1 Pro is ahead overall as of 2026-09-13, but check the category scores for your use.
Is Deepseek v4 Pro better than Gemini 3.1 Pro for coding?
Gemini 3.1 Pro scores higher in coding (60 vs 59 on the category index, where 50 is average).