BenchLeader

GLM-5.2 vs GPT-5.4

Verdict
  • GPT-5.4 (xhigh) leads on quality: 65.4 vs 63.9.
  • GLM-5.2 (max) is stronger in agents & tools, composite, human preference, reasoning.
  • GPT-5.4 (xhigh) is stronger in coding, instruction following, knowledge, long context, maths, multimodal.
  • GLM-5.2 (max) is 2.6× cheaper ($2.15 vs $5.63 per 1M blended).
  • GPT-5.4 (xhigh) streams 1.7× faster (121 vs 71 tokens per second).
MetricGLM-5.2 (max)GPT-5.4 (xhigh)
BenchLeader Index63.965.4
Agents & tools score63.863.4
Coding score63.866.6
Composite score72.169.4
Human preference score67.2
Instruction following score71.672.1
Knowledge score55.560.6
Long context score65.667.5
Maths score59.066.4
Reasoning score72.972.4
Multimodal score62.9
Blended price $/M$2.15$5.63
Output speed71 tok/s121 tok/s
Time to first answer31.6 s141.3 s
Context window1M1.1M
GPQA Diamond91.9%93.3%
FrontierMath Tiers 1–359.2%78.6%
FrontierMath Tier 429.3%49.0%
OTIS Mock AIME86.4%95.3%
SWE-bench Verified (Epoch)78.7%
SimpleQA Verified34.2%45.1%
Humanity's Last Exam36.2%
SciCode50.5%56.6%
WeirdML70.1%77.7%
APEX-Agents36.0%
ProofBench35.0%56.0%
GSO-Bench31.4%
LMArena Text1472
LMArena Hard Prompts1493
LMArena Coding1510
LMArena WebDev1592
LMArena Agent4.4
LiveBench78.0%
LiveBench Reasoning88.1%
LiveBench Coding77.5%
LiveBench Agentic Coding53.8%
LiveBench Mathematics94.2%
LiveBench Data Analysis79.3%
LiveBench Language82.6%
LiveBench Instruction Following70.2%
AA Intelligence Index34.039.0
IFBench73.3%74.0%
AA-LCR78.3%82.0%
MMMU-Pro78.4%
AA-Omniscience4.45.8
Terminal-Bench Hard50.8%57.6%
GPQA Diamond (AA)89.5%92.0%
Humanity's Last Exam (AA)41.1%43.7%
SciCode (AA)51.2%
τ²-Bench Telecom (AA)99.1%87.1%
AIME (Vals)96.7%
LiveCodeBench84.1%
MMLU-Pro87.5%
LegalBench86.0%
CorpFin65.3%
TaxEval74.0%
MedQA96.1%
Terminal-Bench 2.1 (Vals)67.8%
SWE-bench (Vals)82.8%78.2%
GPQA Diamond (Vals)91.7%
AIME 202699.2%
HMMT February 202697.7%
MathArena Apex54.2%
SWE-Bench Pro59.1%
MCP Atlas70.6%
ARC-AGI-193.7%
ARC-AGI-274.0%
CritPt20.9%23.4%
GDPval (AA)45.3%40.3%
τ²-Bench Banking (AA)34.6%39.6%
ITBench SRE (AA)42.7%
APEX-Agents (AA)33.7%33.3%
CaseLaw v263.8%
Code Migration37.9%35.0%
Harvey's Legal Agent Benchmark7.1%0.0%
Legal Research Bench31.3%
MedCode41.3%
MedScribe77.5%
MMMU-Pro (Vals)87.5%
MortgageTax68.3%
ProgramBench0.5%0.0%
SAGE43.3%
SkillsBench45.1%51.7%
SREBench0.0%
Vibe Code Bench v1.164.0%67.4%
SWE-Bench Pro (private)43.4%
SWE Atlas: Codebase QnA40.8%
SWE Atlas: Refactoring44.3%
SWE Atlas: Test Writing44.4%
LMArena Maths1480
LMArena Creative Writing1451
LMArena Instruction Following1466
LMArena Multi-turn1469
LMArena Longer Queries1483
Chess Puzzles21.0%44.0%
EBR-bench9.5%25.4%
Mystery Game Puzzles37.0%
PostTrainBench31.7%
CL-bench27.9%
CL-bench Life21.7%
METR Time Horizons74.3%
DeepSWE43.8%51.8%
LMCA45.8%52.0%
DTBench93.6%94.4%
CursorBench55.0%
ALE-Bench1010.2

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

GLM-5.2 vs GPT-5.4: questions

Is GLM-5.2 better than GPT-5.4?
GPT-5.4 (xhigh) leads on quality: 65.4 vs 63.9. The BenchLeader Index combines every independent quality benchmark; GPT-5.4 (xhigh) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is GLM-5.2 better than GPT-5.4 for coding?
GPT-5.4 scores higher in coding (67 vs 64 on the category index, where 50 is average).
Is GLM-5.2 better than GPT-5.4 for agentic tasks?
GLM-5.2 scores higher in agentic tasks (64 vs 63 on the category index, where 50 is average).
Which is cheaper, GLM-5.2 or GPT-5.4?
GLM-5.2 is cheaper: $2.15 against $5.63 per million tokens, blended at three input tokens per output token.
Which is faster, GLM-5.2 or GPT-5.4?
GPT-5.4 streams faster: 121 against 71 output tokens per second.
Which has the larger context window?
GPT-5.4 accepts more context: 1.1M against 1M tokens.