BenchLeader

GLM-5.3 vs GPT-5.2-Codex

Verdict
  • GLM-5.3 and GPT-5.2-Codex are level on quality (61.4 vs 60.8).
  • GLM-5.3 is stronger in agents & tools, composite, knowledge.
  • GPT-5.2-Codex is stronger in long context, coding, instruction following, multimodal.
  • GLM-5.3 is 2.2× cheaper ($2.15 vs $4.81 per 1M blended).
  • GLM-5.3 streams 2.2× faster (67 vs 30 tokens per second).
MetricGLM-5.3GPT-5.2-Codex
BenchLeader Index61.460.8
Agents & tools score65.263.5
Composite score70.456.8
Knowledge score71.063.2
Long context score66.367.7
Coding score55.7
Instruction following score75.1
Multimodal score60.7
Blended price $/M$2.15$4.81
Output speed67 tok/s30 tok/s
Time to first answer33.0 s3.4 s
Context window1M400k
Terminal-Bench66.5%
APEX-Agents27.6%
Epoch Capabilities Index155.3
LMArena WebDev1339
LiveBench76.1%74.0%
LiveBench Reasoning85.8%77.7%
LiveBench Coding79.0%83.6%
LiveBench Agentic Coding60.9%49.4%
LiveBench Mathematics87.9%88.8%
LiveBench Data Analysis70.2%78.2%
LiveBench Language79.9%73.7%
AA Intelligence Index44.928.5
IFBench77.6%
AA-LCR79.7%82.3%
MMMU-Pro76.3%
AA-Omniscience14.3-2.2
Terminal-Bench Hard37.1%
GPQA Diamond (AA)91.7%89.9%
Humanity's Last Exam (AA)42.3%35.7%
SciCode (AA)59.0%
τ²-Bench Telecom (AA)92.1%
LiveCodeBench88.0%
SWE-Bench Pro41.0%
MCP Atlas84.2%
SWE-bench Verified (bash only)72.8%
SWE-bench Verified (any scaffold)72.8%

Data as of 2026-09-17. Best configuration of each model; every score links to its source on the model pages.

GLM-5.3 vs GPT-5.2-Codex: questions

Is GLM-5.3 better than GPT-5.2-Codex?
GLM-5.3 and GPT-5.2-Codex are level on quality (61.4 vs 60.8). The BenchLeader Index combines every independent quality benchmark; GLM-5.3 is ahead overall as of 2026-09-17, but check the category scores for your use.
Is GLM-5.3 better than GPT-5.2-Codex for agentic tasks?
GLM-5.3 scores higher in agentic tasks (65 vs 64 on the category index, where 50 is average).
Which is cheaper, GLM-5.3 or GPT-5.2-Codex?
GLM-5.3 is cheaper: $2.15 against $4.81 per million tokens, blended at three input tokens per output token.
Which is faster, GLM-5.3 or GPT-5.2-Codex?
GLM-5.3 streams faster: 67 against 30 output tokens per second.
Which has the larger context window?
GLM-5.3 accepts more context: 1M against 400k tokens.