BenchLeader

Gemini 3.1 Pro vs GLM-5.2

Verdict
  • Gemini 3.1 Pro and GLM-5.2 (max) are level on quality (63.9 vs 63.9).
  • Gemini 3.1 Pro is stronger in knowledge, long context, maths, multimodal.
  • GLM-5.2 (max) is stronger in agents & tools, coding, composite, human preference, instruction following, reasoning.
  • GLM-5.2 (max) is 2.1× cheaper ($2.15 vs $4.50 per 1M blended).
  • Gemini 3.1 Pro streams 1.7× faster (121 vs 71 tokens per second).
MetricGemini 3.1 ProGLM-5.2 (max)
BenchLeader Index63.963.9
Agents & tools score60.463.8
Coding score59.863.8
Composite score67.572.1
Human preference score59.867.2
Instruction following score69.771.6
Knowledge score65.955.5
Long context score67.565.6
Maths score61.959.0
Multimodal score66.0
Reasoning score70.872.9
Blended price $/M$4.50$2.15
Output speed121 tok/s71 tok/s
Time to first answer34.8 s31.6 s
Context window1.0M1M
GPQA Diamond94.1%91.9%
FrontierMath Tiers 1–359.6%59.2%
FrontierMath Tier 426.8%29.3%
OTIS Mock AIME95.6%86.4%
SWE-bench Verified (Epoch)78.7%
SimpleQA Verified34.2%
Humanity's Last Exam46.4%
Terminal-Bench80.2%
SimpleBench79.6%
SciCode58.9%50.5%
WeirdML72.1%70.1%
APEX-Agents33.5%
ProofBench26.0%35.0%
GSO-Bench22.6%
Epoch Capabilities Index155
LMArena Text14871472
LMArena Hard Prompts15081493
LMArena Coding15201510
LMArena WebDev14471592
LMArena Vision1296
LMArena Agent-5.84.4
AA Intelligence Index30.434.0
IFBench77.1%73.3%
AA-LCR82.0%78.3%
MMMU-Pro82.4%
AA-Omniscience31.94.4
Terminal-Bench Hard53.8%50.8%
GPQA Diamond (AA)94.1%89.5%
Humanity's Last Exam (AA)47.0%41.1%
SciCode (AA)58.7%51.2%
τ²-Bench Telecom (AA)95.6%99.1%
Terminal-Bench 2.1 (Vals)67.8%
SWE-bench (Vals)82.8%
AIME 202698.3%
HMMT February 202694.7%
MathArena Apex60.9%
MultiChallenge71.4%
PRBench Finance41.9%
PRBench Legal44.0%
MultiNRC64.7%
HiL-Bench35.3%
TutorBench53.0%
EQ-Bench 41142
ARC-AGI-198.0%
ARC-AGI-277.1%
ARC-AGI-30.4%
CritPt17.7%20.9%
GDPval (AA)20.2%45.3%
τ²-Bench Banking (AA)21.4%34.6%
ITBench SRE (AA)30.3%42.7%
Analyst Agent (AA)41.3%
APEX-Agents (AA)32.0%33.7%
Code Migration37.9%
CyberBench36.4%
Harvey's Legal Agent Benchmark7.1%
Legal Research Bench31.3%
ProgramBench0.5%
SkillsBench45.1%
SREBench0.0%
Vibe Code Bench v1.164.0%
FORTRESS29.8%
MASK42.4%
SWE Atlas: Codebase QnA13.5%
SWE Atlas: Refactoring33.8%
SWE Atlas: Test Writing29.8%
VTB29.0%
LMArena Maths14891480
LMArena Creative Writing14801451
LMArena Instruction Following14811466
LMArena Multi-turn14951469
LMArena Longer Queries15001483
LMArena Document1459
Chess Puzzles55.0%21.0%
EBR-bench14.3%9.5%
BALROG57.0%
PostTrainBench22.0%31.7%
ExploitBench26.1%
CL-bench20.8%
CL-bench Life16.9%
METR Time Horizons77.0%
DeepSWE11.7%43.8%
LMCA45.8%
DTBench93.6%
ForecastBench59.0%
CursorBench55.0%
GBAEval0.8%
ALE-Bench1160.61010.2
AlgoTune2.0
Vending-Bench 2911.2
Blueprint-Bench 226.5%
GDP.pdf17.0%

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

Gemini 3.1 Pro vs GLM-5.2: questions

Is Gemini 3.1 Pro better than GLM-5.2?
Gemini 3.1 Pro and GLM-5.2 (max) are level on quality (63.9 vs 63.9). The BenchLeader Index combines every independent quality benchmark; GLM-5.2 (max) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is Gemini 3.1 Pro better than GLM-5.2 for coding?
GLM-5.2 scores higher in coding (64 vs 60 on the category index, where 50 is average).
Is Gemini 3.1 Pro better than GLM-5.2 for agentic tasks?
GLM-5.2 scores higher in agentic tasks (64 vs 60 on the category index, where 50 is average).
Which is cheaper, Gemini 3.1 Pro or GLM-5.2?
GLM-5.2 is cheaper: $2.15 against $4.50 per million tokens, blended at three input tokens per output token.
Which is faster, Gemini 3.1 Pro or GLM-5.2?
Gemini 3.1 Pro streams faster: 121 against 71 output tokens per second.
Which has the larger context window?
Gemini 3.1 Pro accepts more context: 1.0M against 1M tokens.