BenchLeader

GLM-5.2 vs GPT-5.6 Terra

Verdict
  • GPT-5.6 Terra (max) leads on quality: 65.0 vs 63.9.
  • GLM-5.2 (max) is stronger in agents & tools, composite, human preference, instruction following, reasoning.
  • GPT-5.6 Terra (max) is stronger in coding, knowledge, long context, maths, multimodal.
  • GLM-5.2 (max) is 2.1× cheaper ($2.15 vs $4.50 per 1M blended).
MetricGLM-5.2 (max)GPT-5.6 Terra (max)
BenchLeader Index63.965.0
Agents & tools score63.859.0
Coding score63.867.2
Composite score72.171.5
Human preference score67.2
Instruction following score71.669.8
Knowledge score55.558.8
Long context score65.668.0
Maths score59.071.5
Reasoning score72.972.3
Multimodal score65.3
Blended price $/M$2.15$4.50
Output speed71 tok/s84 tok/s
Time to first answer31.6 s220.9 s
Context window1M1.1M
GPQA Diamond91.9%93.3%
FrontierMath Tiers 1–359.2%86.0%
FrontierMath Tier 429.3%70.7%
OTIS Mock AIME86.4%99.7%
SWE-bench Verified (Epoch)78.7%
SimpleQA Verified34.2%43.2%
Terminal-Bench21.5%
SciCode50.5%53.9%
WeirdML70.1%
ProofBench35.0%
LMArena Text1472
LMArena Hard Prompts1493
LMArena Coding1510
LMArena WebDev1592
LMArena Agent4.4
LiveBench77.9%
LiveBench Reasoning90.6%
LiveBench Coding78.3%
LiveBench Agentic Coding55.0%
LiveBench Mathematics94.9%
LiveBench Data Analysis79.3%
LiveBench Language82.9%
LiveBench Instruction Following64.6%
AA Intelligence Index34.042.3
IFBench73.3%71.2%
AA-LCR78.3%83.0%
MMMU-Pro80.7%
AA-Omniscience4.40.1
Terminal-Bench Hard50.8%57.6%
GPQA Diamond (AA)89.5%92.5%
Humanity's Last Exam (AA)41.1%42.9%
SciCode (AA)51.2%55.0%
τ²-Bench Telecom (AA)99.1%86.3%
IOI87.6%
LegalBench85.1%
Terminal-Bench 2.1 (Vals)67.8%77.5%
SWE-bench (Vals)82.8%95.4%
Vals Index59.6
ARC-AGI-196.5%
ARC-AGI-283.9%
ARC-AGI-30.8%
CritPt20.9%30.0%
GDPval (AA)45.3%48.9%
τ²-Bench Banking (AA)34.6%40.2%
ITBench SRE (AA)42.7%51.0%
APEX-Agents (AA)33.7%38.9%
Code Migration37.9%47.8%
Excel Modeling Benchmark66.2%
Finance Agent v254.4%
Harvey's Legal Agent Benchmark7.1%0.8%
Legal Research Bench31.3%41.4%
MMMU-Pro (Vals)86.5%
ProgramBench0.5%0.5%
SAGE47.0%
SkillsBench45.1%58.9%
SREBench0.0%
Tax Agent Bench65.2%
Terminal-Bench 4.0 (Vals)26.3%
Terminal-Bench Science10.0%
Vibe Code Bench 1-10014.8%
Vibe Code Bench v1.164.0%74.6%
LMArena Maths1480
LMArena Creative Writing1451
LMArena Instruction Following1466
LMArena Multi-turn1469
LMArena Longer Queries1483
Chess Puzzles21.0%54.0%
EBR-bench9.5%
Mystery Game Puzzles35.0%
PostTrainBench31.7%
DeepSWE43.8%69.6%
LMCA45.8%52.2%
DTBench93.6%92.5%
CursorBench55.0%64.9%
ALE-Bench1010.21951.4

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

GLM-5.2 vs GPT-5.6 Terra: questions

Is GLM-5.2 better than GPT-5.6 Terra?
GPT-5.6 Terra (max) leads on quality: 65.0 vs 63.9. The BenchLeader Index combines every independent quality benchmark; GPT-5.6 Terra (max) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is GLM-5.2 better than GPT-5.6 Terra for coding?
GPT-5.6 Terra scores higher in coding (67 vs 64 on the category index, where 50 is average).
Is GLM-5.2 better than GPT-5.6 Terra for agentic tasks?
GLM-5.2 scores higher in agentic tasks (64 vs 59 on the category index, where 50 is average).
Which is cheaper, GLM-5.2 or GPT-5.6 Terra?
GLM-5.2 is cheaper: $2.15 against $4.50 per million tokens, blended at three input tokens per output token.
Which is faster, GLM-5.2 or GPT-5.6 Terra?
GPT-5.6 Terra streams faster: 84 against 71 output tokens per second.
Which has the larger context window?
GPT-5.6 Terra accepts more context: 1.1M against 1M tokens.