BenchLeader

Gemini 3.1 Pro vs GPT-5.2-Codex

Verdict
  • Gemini 3.1 Pro leads on quality: 64.0 vs 60.8.
  • Gemini 3.1 Pro is stronger in agents & tools, coding, composite, human preference, knowledge, maths, multimodal, reasoning.
  • GPT-5.2-Codex is stronger in instruction following, long context.
  • They cost about the same ($4.50 per 1M blended).
  • Gemini 3.1 Pro streams 4.3× faster (108 vs 25 tokens per second).
MetricGemini 3.1 ProGPT-5.2-Codex
BenchLeader Index64.060.8
Agents & tools score63.763.5
Coding score59.955.7
Composite score67.456.8
Human preference score59.9
Instruction following score69.675.1
Knowledge score66.263.1
Long context score67.667.8
Maths score60.7
Multimodal score66.260.7
Reasoning score69.8
Blended price $/M$4.50$4.81
Output speed108 tok/s25 tok/s
Time to first answer30.5 s6.7 s
Context window1.0M400k
GPQA Diamond94.1%
FrontierMath Tiers 1–359.6%
FrontierMath Tier 426.8%
OTIS Mock AIME95.6%
Humanity's Last Exam46.4%
Terminal-Bench80.2%66.5%
SimpleBench79.6%
SciCode58.9%
WeirdML72.1%
APEX-Agents33.5%27.6%
ProofBench26.0%
GSO-Bench22.6%
Epoch Capabilities Index155
LMArena Text1487
LMArena Hard Prompts1507
LMArena Coding1521
LMArena WebDev14471339
LMArena Vision1295
LMArena Agent-5.3
LiveBench74.0%
LiveBench Reasoning77.7%
LiveBench Coding83.6%
LiveBench Agentic Coding49.4%
LiveBench Mathematics88.8%
LiveBench Data Analysis78.2%
LiveBench Language73.7%
AA Intelligence Index30.428.5
IFBench77.1%77.6%
AA-LCR82.0%82.3%
MMMU-Pro82.4%76.3%
AA-Omniscience31.9-2.2
Terminal-Bench Hard53.8%37.1%
GPQA Diamond (AA)94.1%89.9%
Humanity's Last Exam (AA)47.0%35.7%
SciCode (AA)58.7%
τ²-Bench Telecom (AA)95.6%92.1%
LiveCodeBench88.0%
AIME 202698.3%
HMMT February 202694.7%
MathArena Apex60.9%
SWE-Bench Pro41.0%
MultiChallenge71.4%
PRBench Finance41.9%
PRBench Legal44.0%
MultiNRC64.7%
HiL-Bench35.3%
TutorBench53.0%
EQ-Bench 41142
SWE-bench Verified (bash only)72.8%
SWE-bench Verified (any scaffold)72.8%
ARC-AGI-198.0%
ARC-AGI-277.1%
ARC-AGI-30.4%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Gemini 3.1 Pro vs GPT-5.2-Codex: questions

Is Gemini 3.1 Pro better than GPT-5.2-Codex?
Gemini 3.1 Pro leads on quality: 64.0 vs 60.8. The BenchLeader Index combines every independent quality benchmark; Gemini 3.1 Pro is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Gemini 3.1 Pro better than GPT-5.2-Codex for coding?
Gemini 3.1 Pro scores higher in coding (60 vs 56 on the category index, where 50 is average).
Is Gemini 3.1 Pro better than GPT-5.2-Codex for agentic tasks?
Gemini 3.1 Pro scores higher in agentic tasks (64 vs 64 on the category index, where 50 is average).
Which is cheaper, Gemini 3.1 Pro or GPT-5.2-Codex?
Gemini 3.1 Pro is cheaper: $4.50 against $4.81 per million tokens, blended at three input tokens per output token.
Which is faster, Gemini 3.1 Pro or GPT-5.2-Codex?
Gemini 3.1 Pro streams faster: 108 against 25 output tokens per second.
Which has the larger context window?
Gemini 3.1 Pro accepts more context: 1.0M against 400k tokens.