BenchLeader

Gemini 3.1 Pro vs GPT-5.2

Verdict
  • Gemini 3.1 Pro leads on quality: 64.0 vs 61.6.
  • Gemini 3.1 Pro is stronger in agents & tools, coding, instruction following, knowledge, maths, multimodal, reasoning.
  • GPT-5.2 is stronger in composite, human preference, long context.
  • They cost about the same ($4.50 per 1M blended).
  • Gemini 3.1 Pro streams 1.6× faster (108 vs 68 tokens per second).
MetricGemini 3.1 ProGPT-5.2
BenchLeader Index64.061.6
Agents & tools score63.761.2
Coding score59.953.7
Composite score67.467.5
Human preference score59.967.9
Instruction following score69.667.4
Knowledge score66.261.1
Long context score67.667.9
Maths score60.7
Multimodal score66.259.9
Reasoning score69.861.3
Blended price $/M$4.50$4.81
Output speed108 tok/s68 tok/s
Time to first answer30.5 s129.4 s
Context window1.0M400k
GPQA Diamond94.1%
FrontierMath Tiers 1–359.6%
FrontierMath Tier 426.8%
OTIS Mock AIME95.6%
Humanity's Last Exam46.4%27.8%
Terminal-Bench80.2%64.9%
SimpleBench79.6%45.8%
SciCode58.9%
Remote Labor Index2.1%
WeirdML72.1%
APEX-Agents33.5%23.0%
ProofBench26.0%
GSO-Bench22.6%
Epoch Capabilities Index155153.5
LMArena Text14871476
LMArena Hard Prompts15071497
LMArena Coding15211515
LMArena WebDev14471417
LMArena Vision12951268
LMArena Agent-5.3
AA Intelligence Index30.430.4
IFBench77.1%75.4%
AA-LCR82.0%82.7%
MMMU-Pro82.4%
AA-Omniscience31.9-0.9
Terminal-Bench Hard53.8%47.0%
GPQA Diamond (AA)94.1%90.3%
Humanity's Last Exam (AA)47.0%37.7%
SciCode (AA)58.7%
τ²-Bench Telecom (AA)95.6%84.8%
AIME 202698.3%
HMMT February 202694.7%
MathArena Apex60.9%
SWE-Bench Pro29.9%
MultiChallenge71.4%
PRBench Finance41.9%
PRBench Legal44.0%
VISTA46.6%
MultiNRC64.7%42.2%
HiL-Bench35.3%
TutorBench53.0%53.5%
EQ-Bench 41142
Kagi LLM Benchmark73.3%
SWE-bench Verified (bash only)69.0%
ARC-AGI-198.0%94.5%
ARC-AGI-277.1%72.9%
ARC-AGI-30.4%
BFCL Overall55.9%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Gemini 3.1 Pro vs GPT-5.2: questions

Is Gemini 3.1 Pro better than GPT-5.2?
Gemini 3.1 Pro leads on quality: 64.0 vs 61.6. The BenchLeader Index combines every independent quality benchmark; Gemini 3.1 Pro is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Gemini 3.1 Pro better than GPT-5.2 for coding?
Gemini 3.1 Pro scores higher in coding (60 vs 54 on the category index, where 50 is average).
Is Gemini 3.1 Pro better than GPT-5.2 for agentic tasks?
Gemini 3.1 Pro scores higher in agentic tasks (64 vs 61 on the category index, where 50 is average).
Which is cheaper, Gemini 3.1 Pro or GPT-5.2?
Gemini 3.1 Pro is cheaper: $4.50 against $4.81 per million tokens, blended at three input tokens per output token.
Which is faster, Gemini 3.1 Pro or GPT-5.2?
Gemini 3.1 Pro streams faster: 108 against 68 output tokens per second.
Which has the larger context window?
Gemini 3.1 Pro accepts more context: 1.0M against 400k tokens.