BenchLeader

Deepseek v4 Pro vs Gemini 3.5 Flash

Verdict
  • Gemini 3.5 Flash (medium) leads on quality: 62.2 vs 60.4.
  • Deepseek v4 Pro (high) is stronger in coding, maths.
  • Gemini 3.5 Flash (medium) is stronger in human preference, reasoning, agents & tools, composite, instruction following, knowledge, long context, multimodal.
MetricDeepseek v4 Pro (high)Gemini 3.5 Flash (medium)
BenchLeader Index60.462.2
Coding score59.358.7
Human preference score66.067.5
Maths score66.2
Reasoning score65.967.0
Agents & tools score63.3
Composite score65.3
Instruction following score69.4
Knowledge score70.7
Long context score60.7
Multimodal score67.3
Blended price $/M$3.38
Output speed234 tok/s
Time to first answer11.2 s
Context window1.0M
GPQA Diamond90.9%
OTIS Mock AIME95.6%
SciCode46.4%
WeirdML46.5%
LMArena Text14621474
LMArena Hard Prompts14821493
LMArena Coding15051502
LMArena WebDev15811492
LMArena Vision1306
AA Intelligence Index33.6
IFBench74.6%
AA-LCR74.3%
MMMU-Pro83.9%
AA-Omniscience20.8
Terminal-Bench Hard39.4%
GPQA Diamond (AA)92.1%
Humanity's Last Exam (AA)41.3%
τ²-Bench Telecom (AA)95.6%

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

Deepseek v4 Pro vs Gemini 3.5 Flash: questions

Is Deepseek v4 Pro better than Gemini 3.5 Flash?
Gemini 3.5 Flash (medium) leads on quality: 62.2 vs 60.4. The BenchLeader Index combines every independent quality benchmark; Gemini 3.5 Flash (medium) is ahead overall as of 2026-09-13, but check the category scores for your use.
Is Deepseek v4 Pro better than Gemini 3.5 Flash for coding?
Deepseek v4 Pro scores higher in coding (59 vs 59 on the category index, where 50 is average).