BenchLeader

Gemini 3 Pro vs Gemini 4 Argon

Verdict
  • Gemini 4 Argon (high) leads on quality: 70.0 vs 60.6.
  • Gemini 3 Pro is stronger in instruction following, multimodal.
  • Gemini 4 Argon (high) is stronger in agents & tools, coding, human preference, knowledge, maths, reasoning, composite, long context.
MetricGemini 3 ProGemini 4 Argon (high)
BenchLeader Index60.670.0
Agents & tools score61.867.5
Coding score57.972.1
Human preference score68.373.2
Instruction following score61.0–
Knowledge score56.475.7
Maths score54.471.9
Multimodal score65.6–
Reasoning score65.882.0
Composite score–89.2
Long context score–65.0
Blended price $/M–$4.00
Context window–1M
GPQA Diamond92.6%–
OTIS Mock AIME91.4%–
SWE-bench Verified (Epoch)72.9%–
Humanity's Last Exam37.5%–
Terminal-Bench69.4%–
SimpleBench76.4%–
GDPval40.3%–
SciCode–61.8%
Remote Labor Index1.3%–
WeirdML69.9%–
ProofBench20.0%–
GSO-Bench18.6%–
Epoch Capabilities Index152.9–
LMArena Text14861525
LMArena Hard Prompts15021551
LMArena Coding15171562
LMArena WebDev14401678
LMArena Vision1305–
LMArena Agent–9.3
AA Intelligence Index v4.3.2–52.6
AA-LCR–79.7%
AA-Omniscience–42.4
Humanity's Last Exam (AA)–57.1%
SciCode (AA)–61.8%
LiveCodeBench86.4%–
MMLU-Pro90.1%–
IOI–100.0%
LegalBench–88.3%
Vals Index–68.9
AIME 202691.7%–
HMMT February 202686.4%–
MathArena Apex23.4%–
SWE-Bench Pro43.3%–
MCP Atlas70.3%–
MultiChallenge65.7%–
PRBench Finance39.2%–
PRBench Legal40.6%–
VISTA51.5%–
MultiNRC59.0%–
TutorBench53.7%–
MMMU-Pro (official)81.0%–
Kagi LLM Benchmark80.1%–
IFEval (HELM)87.6%–
Omni-MATH (HELM)55.6%–
WildBench (HELM)85.9%–
MMLU-Pro (HELM)90.3%–
GPQA Diamond (HELM)80.3%–
HELM Capabilities mean79.9%–
SWE-bench Verified (bash only)74.2%–
SWE-bench Verified (any scaffold)74.2%–
ARC-AGI-175.0%–
ARC-AGI-254.0%–
BFCL Overall72.5%–
CritPt–27.1%
GDPval-AA v2.1–56.3%
BioMysteryBench–76.3%
Code Migration–68.2%
CUA-bench–4.8%
CyberBench–77.9%
Excel Modeling Benchmark–75.2%
Finance Agent v2–65.4%
Harvey's Legal Agent Benchmark–19.6%
Legal Research Bench–54.8%
MedCode–58.8%
MedScribe–87.4%
MysteryMechanism–45.5%
Poker Agent1078.9%–
ProgramBench–2.5%
Public Benefits Bench–69.8%
SAGE–53.6%
SREBench–44.3%
Tax Agent Bench–76.2%
Terminal-Bench 4.0 (Vals)–57.6%
Terminal-Bench Science–44.3%
Vibe Code Bench v1.1–91.9%
FORTRESS41.7%–
MASK42.6%–
PropensityBench52.9%–
SciPredict25.3%–
SWE-Bench Pro (private)43.3%–
LMArena Maths14771528
LMArena Creative Writing14851519
LMArena Instruction Following14731529
LMArena Multi-turn14961553
LMArena Longer Queries14911545
LMArena Document1451–
Chess Puzzles31.0%–
BALROG58.1%–
GeoBench84.0%–
VPCT91.0%–
CL-bench15.8%–
METR Time Horizons71.0%–
ForecastBench61.2%–
ALE-Bench1176.8–
AlgoTune1.8–
Vending-Bench 25478.2–
FrontierSWE–55.0%
Terminal-Bench 4.0 (AA)–57.1%
AutomationBench–77.5%
GDP.pdf–21.8%
AA-Omniscience: accuracy–49.9%
AA-Omniscience: non-hallucination–84.9%
AA-Briefcase v1.1–1488

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Gemini 3 Pro vs Gemini 4 Argon: questions

Is Gemini 3 Pro better than Gemini 4 Argon?
Gemini 4 Argon (high) leads on quality: 70.0 vs 60.6. The BenchLeader Index combines every independent quality benchmark; Gemini 4 Argon (high) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Gemini 3 Pro better than Gemini 4 Argon for coding?
Gemini 4 Argon scores higher in coding (72 vs 58 on the category index, where 50 is average).
Is Gemini 3 Pro better than Gemini 4 Argon for agentic tasks?
Gemini 4 Argon scores higher in agentic tasks (68 vs 62 on the category index, where 50 is average).