BenchLeader

Gemini 3.1 Pro vs o3

Verdict
  • Gemini 3.1 Pro leads on quality: 63.9 vs 61.6.
  • Gemini 3.1 Pro is stronger in composite, instruction following, knowledge, long context, multimodal, reasoning.
  • o3 is stronger in agents & tools, coding, human preference, maths.
  • o3 is 1.3× cheaper ($3.50 vs $4.50 per 1M blended).
MetricGemini 3.1 Proo3
BenchLeader Index63.961.6
Agents & tools score63.470.3
Coding score59.862.0
Composite score67.454.5
Human preference score59.862.3
Instruction following score69.665.0
Knowledge score66.358.0
Long context score67.663.8
Maths score60.778.6
Multimodal score66.257.9
Reasoning score69.861.5
Blended price $/M$4.50$3.50
Output speed108 tok/s107 tok/s
Time to first answer28.3 s6.3 s
Context window1.0M200k
GPQA Diamond94.1%
FrontierMath Tiers 1–359.6%
FrontierMath Tier 426.8%
OTIS Mock AIME95.6%
Humanity's Last Exam46.4%
Terminal-Bench80.2%
SimpleBench79.6%
SciCode58.9%
WeirdML72.1%
APEX-Agents33.5%
ProofBench26.0%
GSO-Bench22.6%
Epoch Capabilities Index155146.9
LMArena Text14871432
LMArena Hard Prompts15081441
LMArena Coding15201460
LMArena WebDev1447
LMArena Vision12951214
LMArena Agent-5.6
AA Intelligence Index30.420.2
IFBench77.1%71.4%
AA-LCR82.0%74.7%
MMMU-Pro82.4%70.1%
AA-Omniscience31.9-15.6
Terminal-Bench Hard53.8%37.1%
GPQA Diamond (AA)94.1%82.7%
Humanity's Last Exam (AA)47.0%20.1%
SciCode (AA)58.7%
τ²-Bench Telecom (AA)95.6%80.7%
AIME 202698.3%
HMMT February 202694.7%
MathArena Apex60.9%
MultiChallenge71.4%
PRBench Finance41.9%47.7%
PRBench Legal44.0%48.6%
MultiNRC64.7%
HiL-Bench35.3%
TutorBench53.0%
EQ-Bench 41142
MMMU (validation)82.9%
MMMU-Pro (official)76.4%
Kagi LLM Benchmark67.6%
IFEval (HELM)86.9%
Omni-MATH (HELM)71.4%
WildBench (HELM)86.1%
MMLU-Pro (HELM)85.9%
GPQA Diamond (HELM)75.3%
HELM Capabilities mean81.1%
Aider Polyglot81.3%
SWE-bench Verified (bash only)58.4%
SWE-bench Verified (any scaffold)58.4%
ARC-AGI-198.0%
ARC-AGI-277.1%
ARC-AGI-30.4%
BFCL Overall63.0%

Data as of 2026-09-12. Best configuration of each model; every score links to its source on the model pages.

Gemini 3.1 Pro vs o3: questions

Is Gemini 3.1 Pro better than o3?
Gemini 3.1 Pro leads on quality: 63.9 vs 61.6. The BenchLeader Index combines every independent quality benchmark; Gemini 3.1 Pro is ahead overall as of 2026-09-12, but check the category scores for your use.
Is Gemini 3.1 Pro better than o3 for coding?
o3 scores higher in coding (62 vs 60 on the category index, where 50 is average).
Is Gemini 3.1 Pro better than o3 for agentic tasks?
o3 scores higher in agentic tasks (70 vs 63 on the category index, where 50 is average).
Which is cheaper, Gemini 3.1 Pro or o3?
o3 is cheaper: $3.50 against $4.50 per million tokens, blended at three input tokens per output token.
Which is faster, Gemini 3.1 Pro or o3?
Gemini 3.1 Pro streams faster: 108 against 107 output tokens per second.
Which has the larger context window?
Gemini 3.1 Pro accepts more context: 1.0M against 200k tokens.