BenchLeader

Claude Opus 5 vs Gemini 3 Pro

Verdict
  • Claude Opus 5 leads on quality: 70.0 vs 61.7.
  • Claude Opus 5 is stronger in agents & tools, coding, composite, human preference, knowledge, long context, maths, multimodal, reasoning.
  • Gemini 3 Pro is stronger in instruction following.
  • Gemini 3 Pro is 2.2× cheaper ($4.50 vs $10.00 per 1M blended).
MetricClaude Opus 5Gemini 3 Pro
BenchLeader Index70.061.7
Agents & tools score71.562.3
Coding score74.558.9
Composite score93.164.3
Human preference score78.269.0
Knowledge score70.660.9
Long context score66.264.5
Maths score67.148.4
Multimodal score69.465.7
Reasoning score71.567.1
Instruction following score64.3
Blended price $/M$10.00$4.50
Output speed51 tok/s
Time to first answer93.0 s
Context window1M1M
GPQA Diamond92.9%92.6%
OTIS Mock AIME97.8%91.4%
SWE-bench Verified (Epoch)72.9%
Humanity's Last Exam37.5%
Terminal-Bench69.4%
SimpleBench80.6%76.4%
GDPval40.3%
Remote Labor Index1.3%
WeirdML86.3%69.9%
APEX-Agents31.5%
ProofBench20.0%
GSO-Bench18.6%
Epoch Capabilities Index162.6153.0
LMArena Text1486
LMArena Hard Prompts1504
LMArena Coding1518
LMArena WebDev1439
LMArena Vision1305
AA Intelligence Index50.728.0
IFBench70.4%
AA-LCR79.3%76.0%
MMMU-Pro84.7%80.2%
AA-Omniscience37.115.3
Terminal-Bench Hard41.7%
GPQA Diamond (AA)93.2%90.8%
Humanity's Last Exam (AA)54.9%39.7%
SciCode (AA)56.4%
τ²-Bench Telecom (AA)87.1%
LiveCodeBench89.0%86.4%
MMLU-Pro91.6%90.1%
IOI91.7%
LegalBench87.0%
CorpFin73.2%
TaxEval75.1%
Terminal-Bench 2.1 (Vals)84.6%
SWE-bench (Vals)97.0%
GPQA Diamond (Vals)93.4%
Vals Index67.2
AIME 202691.7%
HMMT February 202686.4%
MathArena Apex23.4%
SWE-Bench Pro43.3%
MCP Atlas70.3%
MultiChallenge65.7%
PRBench Finance39.2%
PRBench Legal40.6%
VISTA51.5%
MultiNRC59.0%
HiL-Bench57.0%
TutorBench53.7%
EQ-Bench 41385
MMMU-Pro (official)81.0%
Kagi LLM Benchmark80.1%
SWE-bench Verified (bash only)74.2%
SWE-bench Verified (any scaffold)77.4%
ARC-AGI-175.0%
ARC-AGI-254.0%
BFCL Overall72.5%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Claude Opus 5 vs Gemini 3 Pro: questions

Is Claude Opus 5 better than Gemini 3 Pro?
Claude Opus 5 leads on quality: 70.0 vs 61.7. The BenchLeader Index combines every independent quality benchmark; Claude Opus 5 is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Claude Opus 5 better than Gemini 3 Pro for coding?
Claude Opus 5 scores higher in coding (75 vs 59 on the category index, where 50 is average).
Is Claude Opus 5 better than Gemini 3 Pro for agentic tasks?
Claude Opus 5 scores higher in agentic tasks (72 vs 62 on the category index, where 50 is average).
Which is cheaper, Claude Opus 5 or Gemini 3 Pro?
Gemini 3 Pro is cheaper: $4.50 against $10.00 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Both accept 1M tokens of context.