BenchLeader

Claude Sonnet 5.5 vs Gemini 3 Pro

Verdict
  • Claude Sonnet 5.5 (max) leads on quality: 67.6 vs 60.6.
  • Claude Sonnet 5.5 (max) is stronger in agents & tools, coding, composite, knowledge, long context, maths, reasoning.
  • Gemini 3 Pro is stronger in human preference, instruction following, multimodal.
MetricClaude Sonnet 5.5 (max)Gemini 3 Pro
BenchLeader Index67.660.6
Agents & tools score69.561.8
Coding score69.857.9
Composite score72.6–
Knowledge score65.456.4
Long context score66.6–
Maths score72.954.4
Reasoning score81.165.8
Human preference score–68.3
Instruction following score–61.0
Multimodal score–65.6
Blended price $/M$4.00–
Output speed141 tok/s–
Time to first answer487.1 s–
Context window1M–
GPQA Diamond95.6%92.6%
FrontierMath Tiers 1–388.8%–
FrontierMath Tier 480.5%–
OTIS Mock AIME100.0%91.4%
SWE-bench Verified (Epoch)–72.9%
SimpleQA Verified46.5%–
Humanity's Last Exam–37.5%
Terminal-Bench61.8%69.4%
SimpleBench–76.4%
GDPval–40.3%
SciCode61.0%–
Remote Labor Index–1.3%
WeirdML–69.9%
APEX-Agents75.5%–
ProofBench100.0%20.0%
GSO-Bench–18.6%
Epoch Capabilities Index–152.9
LMArena Text–1486
LMArena Hard Prompts–1502
LMArena Coding–1517
LMArena WebDev–1440
LMArena Vision–1305
LMArena Agent12–
LiveBench75.7%–
LiveBench Reasoning91.6%–
LiveBench Coding91.4%–
LiveBench Agentic Coding56.3%–
LiveBench Mathematics96.1%–
LiveBench Data Analysis59.5%–
LiveBench Language78.0%–
LiveBench Instruction Following56.8%–
AA Intelligence Index v4.3.256–
AA-LCR82.7%–
AA-Omniscience32.3–
Humanity's Last Exam (AA)55.0%–
SciCode (AA)61.0%–
LiveCodeBench–86.4%
MMLU-Pro–90.1%
AIME 2026–91.7%
HMMT February 2026–86.4%
MathArena Apex–23.4%
SWE-Bench Pro–43.3%
MCP Atlas–70.3%
MultiChallenge–65.7%
PRBench Finance–39.2%
PRBench Legal–40.6%
VISTA–51.5%
MultiNRC–59.0%
TutorBench–53.7%
MMMU-Pro (official)–81.0%
Kagi LLM Benchmark–80.1%
IFEval (HELM)–87.6%
Omni-MATH (HELM)–55.6%
WildBench (HELM)–85.9%
MMLU-Pro (HELM)–90.3%
GPQA Diamond (HELM)–80.3%
HELM Capabilities mean–79.9%
SWE-bench Verified (bash only)–74.2%
SWE-bench Verified (any scaffold)–74.2%
ARC-AGI-1–75.0%
ARC-AGI-2–54.0%
BFCL Overall–72.5%
CritPt31.4%–
GDPval-AA v2.167.0%–
Analyst Agent (AA)57.5%–
Poker Agent–1078.9%
FORTRESS–41.7%
MASK–42.6%
PropensityBench–52.9%
SciPredict–25.3%
SWE-Bench Pro (private)–43.3%
LMArena Maths–1477
LMArena Creative Writing–1485
LMArena Instruction Following–1473
LMArena Multi-turn–1496
LMArena Longer Queries–1491
LMArena Document–1451
Chess Puzzles–31.0%
Mystery Game Puzzles65.0%–
BALROG–58.1%
GeoBench–84.0%
VPCT–91.0%
CL-bench–15.8%
METR Time Horizons–71.0%
ForecastBench–61.2%
CursorBench55.5%–
ALE-Bench–1176.8
AlgoTune–1.8
Vending-Bench 2–5478.2
FrontierSWE61.9%–
Terminal-Bench 4.0 (AA)63.6%–
AutomationBench71.8%–
GDP.pdf25.8%–
MLCR75.0%–
Harvey LAB2.8%–
AA-Omniscience: accuracy54.0%–
AA-Omniscience: non-hallucination53.0%–
AA-Briefcase v1.11823–

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Claude Sonnet 5.5 vs Gemini 3 Pro: questions

Is Claude Sonnet 5.5 better than Gemini 3 Pro?
Claude Sonnet 5.5 (max) leads on quality: 67.6 vs 60.6. The BenchLeader Index combines every independent quality benchmark; Claude Sonnet 5.5 (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Claude Sonnet 5.5 better than Gemini 3 Pro for coding?
Claude Sonnet 5.5 scores higher in coding (70 vs 58 on the category index, where 50 is average).
Is Claude Sonnet 5.5 better than Gemini 3 Pro for agentic tasks?
Claude Sonnet 5.5 scores higher in agentic tasks (70 vs 62 on the category index, where 50 is average).