BenchLeader

Gemini 3.1 Pro vs Gemini 4 Argon

Verdict
  • Gemini 4 Argon (high) leads on quality: 70.0 vs 62.1.
  • Gemini 3.1 Pro is stronger in instruction following, long context, multimodal.
  • Gemini 4 Argon (high) is stronger in agents & tools, coding, composite, human preference, knowledge, maths, reasoning.
  • They cost about the same ($4.00 per 1M blended).
MetricGemini 3.1 ProGemini 4 Argon (high)
BenchLeader Index62.170.0
Agents & tools score56.767.5
Coding score57.372.1
Composite score63.289.2
Human preference score59.573.2
Instruction following score68.4–
Knowledge score64.575.7
Long context score66.365.0
Maths score61.071.9
Multimodal score64.9–
Reasoning score68.682.0
Blended price $/M$4.50$4.00
Output speed114 tok/s–
Time to first answer24.8 s–
Context window1.0M1M
GPQA Diamond94.1%–
FrontierMath Tiers 1–359.6%–
FrontierMath Tier 426.8%–
OTIS Mock AIME95.6%–
Humanity's Last Exam46.4%–
Terminal-Bench80.2%–
SimpleBench79.6%–
SciCode58.9%61.8%
WeirdML72.1%–
APEX-Agents35.3%–
ProofBench26.0%–
GSO-Bench22.6%–
Epoch Capabilities Index154.8–
LMArena Text14871525
LMArena Hard Prompts15081551
LMArena Coding15221562
LMArena WebDev14471678
LMArena Vision1296–
LMArena Agent-7.89.3
AA Intelligence Index v4.3.229.752.6
IFBench77.1%–
AA-LCR82.0%79.7%
MMMU-Pro82.4%–
AA-Omniscience31.942.4
Terminal-Bench Hard53.8%–
GPQA Diamond (AA)94.1%–
Humanity's Last Exam (AA)47.0%57.1%
SciCode (AA)58.7%61.8%
τ²-Bench Telecom (AA)95.6%–
IOI–100.0%
LegalBench–88.3%
Vals Index–68.9
AIME 202698.3%–
HMMT February 202694.7%–
MathArena Apex60.9%–
MultiChallenge71.4%–
PRBench Finance41.9%–
PRBench Legal44.0%–
MultiNRC64.7%–
HiL-Bench35.3%–
TutorBench53.0%–
EQ-Bench 41142–
ARC-AGI-198.0%–
ARC-AGI-277.1%–
ARC-AGI-30.4%–
CritPt17.7%27.1%
GDPval-AA v2.114.7%56.3%
τ³-Banking (AA)21.4%–
ITBench SRE (AA)30.3%–
Analyst Agent (AA)41.3%–
APEX-Agents (AA)32.0%–
BioMysteryBench–76.3%
Code Migration–68.2%
CUA-bench–4.8%
CyberBench–77.9%
Excel Modeling Benchmark–75.2%
Finance Agent v2–65.4%
Harvey's Legal Agent Benchmark–19.6%
Legal Research Bench–54.8%
MedCode–58.8%
MedScribe–87.4%
MysteryMechanism–45.5%
ProgramBench–2.5%
Public Benefits Bench–69.8%
SAGE–53.6%
SREBench–44.3%
Tax Agent Bench–76.2%
Terminal-Bench 4.0 (Vals)–57.6%
Terminal-Bench Science–44.3%
Vibe Code Bench v1.1–91.9%
FORTRESS29.8%–
MASK42.4%–
SWE Atlas: Codebase QnA13.5%–
SWE Atlas: Refactoring33.8%–
SWE Atlas: Test Writing29.8%–
LMArena Maths14881528
LMArena Creative Writing14811519
LMArena Instruction Following14801529
LMArena Multi-turn14971553
LMArena Longer Queries15001545
LMArena Document1459–
Chess Puzzles55.0%–
EBR-bench14.3%–
BALROG57.0%–
PostTrainBench22.0%–
ExploitBench26.1%–
CL-bench20.8%–
CL-bench Life16.9%–
METR Time Horizons77.0%–
DeepSWE v1.111.7%–
ForecastBench59.0%–
GBAEval0.8%–
ALE-Bench1160.6–
AlgoTune2.0–
Vending-Bench 2911.2–
Blueprint-Bench 226.5%–
GDP.pdf17.0%–
FrontierSWE–55.0%
Terminal-Bench 4.0 (AA)4.0%57.1%
Terminal-Bench 2.1 (AA)73.8%–
AutomationBench35.4%77.5%
GDP.pdf17.8%21.8%
MLCR15.6%–
Harvey LAB0.0%–
EnterpriseOps-Gym42.2%–
AA-Omniscience: accuracy54.9%49.9%
AA-Omniscience: non-hallucination49.1%84.9%
AA-Briefcase v1.14561488
BrowseComp31.2%–
DeepSearchQA60.2%–
FACTS Search83.7%–

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Gemini 3.1 Pro vs Gemini 4 Argon: questions

Is Gemini 3.1 Pro better than Gemini 4 Argon?
Gemini 4 Argon (high) leads on quality: 70.0 vs 62.1. The BenchLeader Index combines every independent quality benchmark; Gemini 4 Argon (high) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Gemini 3.1 Pro better than Gemini 4 Argon for coding?
Gemini 4 Argon scores higher in coding (72 vs 57 on the category index, where 50 is average).
Is Gemini 3.1 Pro better than Gemini 4 Argon for agentic tasks?
Gemini 4 Argon scores higher in agentic tasks (68 vs 57 on the category index, where 50 is average).
Which is cheaper, Gemini 3.1 Pro or Gemini 4 Argon?
Gemini 4 Argon is cheaper: $4.00 against $4.50 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Gemini 3.1 Pro accepts more context: 1.0M against 1M tokens.