BenchLeader

Claude Sonnet 5 vs Gemini 3.1 Pro

Verdict
  • Gemini 3.1 Pro leads on quality: 64.0 vs 61.4.
  • Claude Sonnet 5 (high) is stronger in coding, composite, human preference.
  • Gemini 3.1 Pro is stronger in agents & tools, knowledge, long context, multimodal, reasoning, instruction following, maths.
  • They cost about the same ($4.00 per 1M blended).
  • Gemini 3.1 Pro streams 1.8× faster (108 vs 60 tokens per second).
MetricClaude Sonnet 5 (high)Gemini 3.1 Pro
BenchLeader Index61.464.0
Agents & tools score61.863.7
Coding score61.659.9
Composite score69.567.4
Human preference score66.259.9
Knowledge score62.466.2
Long context score64.867.6
Multimodal score63.166.2
Reasoning score66.869.8
Instruction following score69.6
Maths score60.7
Blended price $/M$4.00$4.50
Output speed60 tok/s108 tok/s
Time to first answer10.0 s30.5 s
Context window1M1.0M
GPQA Diamond94.1%
FrontierMath Tiers 1–359.6%
FrontierMath Tier 426.8%
OTIS Mock AIME95.6%
Humanity's Last Exam46.4%
Terminal-Bench80.2%
SimpleBench79.6%
SciCode48.6%58.9%
WeirdML68.8%72.1%
APEX-Agents33.5%
ProofBench26.0%
GSO-Bench22.6%
Epoch Capabilities Index155
LMArena Text14621487
LMArena Hard Prompts14901507
LMArena Coding15211521
LMArena WebDev15381447
LMArena Vision12781295
LMArena Agent6.3-5.3
AA Intelligence Index32.030.4
IFBench77.1%
AA-LCR76.7%82.0%
MMMU-Pro82.4%
AA-Omniscience-3.731.9
Terminal-Bench Hard53.8%
GPQA Diamond (AA)94.1%
Humanity's Last Exam (AA)35.7%47.0%
SciCode (AA)54.3%58.7%
τ²-Bench Telecom (AA)95.6%
AIME 202698.3%
HMMT February 202694.7%
MathArena Apex60.9%
MultiChallenge71.4%
PRBench Finance41.9%
PRBench Legal44.0%
MultiNRC64.7%
HiL-Bench35.3%
TutorBench53.0%
EQ-Bench 41142
ARC-AGI-198.0%
ARC-AGI-277.1%
ARC-AGI-30.4%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Claude Sonnet 5 vs Gemini 3.1 Pro: questions

Is Claude Sonnet 5 better than Gemini 3.1 Pro?
Gemini 3.1 Pro leads on quality: 64.0 vs 61.4. The BenchLeader Index combines every independent quality benchmark; Gemini 3.1 Pro is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Claude Sonnet 5 better than Gemini 3.1 Pro for coding?
Claude Sonnet 5 scores higher in coding (62 vs 60 on the category index, where 50 is average).
Is Claude Sonnet 5 better than Gemini 3.1 Pro for agentic tasks?
Gemini 3.1 Pro scores higher in agentic tasks (64 vs 62 on the category index, where 50 is average).
Which is cheaper, Claude Sonnet 5 or Gemini 3.1 Pro?
Claude Sonnet 5 is cheaper: $4.00 against $4.50 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Sonnet 5 or Gemini 3.1 Pro?
Gemini 3.1 Pro streams faster: 108 against 60 output tokens per second.
Which has the larger context window?
Gemini 3.1 Pro accepts more context: 1.0M against 1M tokens.