BenchLeader

Gemini 3.1 Pro vs GPT-5

Verdict
  • Gemini 3.1 Pro leads on quality: 63.9 vs 60.8.
  • Gemini 3.1 Pro is stronger in agents & tools, composite, instruction following, knowledge, long context, multimodal, reasoning.
  • GPT-5 is stronger in coding, human preference, maths.
  • GPT-5 is 1.3× cheaper ($3.44 vs $4.50 per 1M blended).
MetricGemini 3.1 ProGPT-5
BenchLeader Index63.960.8
Agents & tools score63.451.0
Coding score59.860.9
Composite score67.458.1
Human preference score59.861.7
Instruction following score69.665.0
Knowledge score66.362.2
Long context score67.665.6
Maths score60.772.8
Multimodal score66.259.6
Reasoning score69.864.2
Blended price $/M$4.50$3.44
Output speed108 tok/s85 tok/s
Time to first answer28.3 s60.7 s
Context window1.0M400k
GPQA Diamond94.1%
FrontierMath Tiers 1–359.6%
FrontierMath Tier 426.8%
OTIS Mock AIME95.6%
Humanity's Last Exam46.4%
Terminal-Bench80.2%49.6%
SimpleBench79.6%
SciCode58.9%42.9%
Remote Labor Index1.7%
WeirdML72.1%39.8%
APEX-Agents33.5%18.3%
ProofBench26.0%
GSO-Bench22.6%
Epoch Capabilities Index155150
LMArena Text14871427
LMArena Hard Prompts15081449
LMArena Coding15201462
LMArena WebDev1447
LMArena Vision12951232
LMArena Agent-5.6
AA Intelligence Index30.423.0
IFBench77.1%73.1%
AA-LCR82.0%78.2%
MMMU-Pro82.4%74.2%
AA-Omniscience31.9-8.7
Terminal-Bench Hard53.8%32.6%
GPQA Diamond (AA)94.1%85.3%
Humanity's Last Exam (AA)47.0%28.5%
SciCode (AA)58.7%
τ²-Bench Telecom (AA)95.6%84.8%
AIME 202698.3%
HMMT February 202694.7%
MathArena Apex60.9%
MultiChallenge71.4%
PRBench Finance41.9%51.3%
PRBench Legal44.0%49.0%
VISTA49.7%
MultiNRC64.7%52.1%
HiL-Bench35.3%
TutorBench53.0%55.3%
EQ-Bench 41142
Kagi LLM Benchmark72.7%
IFEval (HELM)87.5%
Omni-MATH (HELM)64.7%
WildBench (HELM)85.7%
MMLU-Pro (HELM)86.3%
GPQA Diamond (HELM)79.1%
HELM Capabilities mean80.7%
Aider Polyglot88.0%
SWE-bench Verified (any scaffold)75.6%
ARC-AGI-198.0%
ARC-AGI-277.1%
ARC-AGI-30.4%

Data as of 2026-09-12. Best configuration of each model; every score links to its source on the model pages.

Gemini 3.1 Pro vs GPT-5: questions

Is Gemini 3.1 Pro better than GPT-5?
Gemini 3.1 Pro leads on quality: 63.9 vs 60.8. The BenchLeader Index combines every independent quality benchmark; Gemini 3.1 Pro is ahead overall as of 2026-09-12, but check the category scores for your use.
Is Gemini 3.1 Pro better than GPT-5 for coding?
GPT-5 scores higher in coding (61 vs 60 on the category index, where 50 is average).
Is Gemini 3.1 Pro better than GPT-5 for agentic tasks?
Gemini 3.1 Pro scores higher in agentic tasks (63 vs 51 on the category index, where 50 is average).
Which is cheaper, Gemini 3.1 Pro or GPT-5?
GPT-5 is cheaper: $3.44 against $4.50 per million tokens, blended at three input tokens per output token.
Which is faster, Gemini 3.1 Pro or GPT-5?
Gemini 3.1 Pro streams faster: 108 against 85 output tokens per second.
Which has the larger context window?
Gemini 3.1 Pro accepts more context: 1.0M against 400k tokens.