BenchLeader

GLM 5.3 vs GLM 5.3 Flash

Verdict
  • GLM 5.3 (max) and GLM 5.3 Flash are level on quality (60.0 vs 60.1).
  • GLM 5.3 (max) is stronger in agents & tools, coding, human preference, maths.
  • GLM 5.3 Flash is stronger in knowledge, reasoning, composite, long context, multimodal.
  • GLM 5.3 Flash is 18× cheaper ($0.119 vs $2.15 per 1M blended).
  • GLM 5.3 Flash streams 1.6× faster (107 vs 66 tokens per second).
MetricGLM 5.3 (max)GLM 5.3 Flash
BenchLeader Index60.060.1
Agents & tools score52.851.5
Coding score64.364.0
Human preference score68.967.6
Knowledge score55.864.3
Maths score60.0
Reasoning score66.967.5
Composite score58.2
Long context score64.1
Multimodal score65.5
Blended price $/M$2.15$0.119
Output speed66 tok/s107 tok/s
Time to first answer33.5 s21.1 s
Context window1M1M
GPQA Diamond90.9%
FrontierMath Tiers 1–368.8%
FrontierMath Tier 429.3%
OTIS Mock AIME91.1%
SimpleQA Verified41.0%
Terminal-Bench41.8%
SciCode56.5%46.1%
WeirdML75.4%
ProofBench49.0%
Epoch Capabilities Index151.4
LMArena Text14861475
LMArena Hard Prompts15091497
LMArena Coding15261524
LMArena WebDev16141607
LMArena Vision1296
LMArena Agent2.61.9
LiveBench71.6%
LiveBench Reasoning77.6%
LiveBench Coding79.0%
LiveBench Agentic Coding56.8%
LiveBench Mathematics81.2%
LiveBench Data Analysis76.4%
LiveBench Language77.3%
AA Intelligence Index41.9
AA-LCR80.0%
AA-Omniscience7.5
GPQA Diamond (AA)91.2%
Humanity's Last Exam (AA)39.9%
SciCode (AA)51.6%
LiveCodeBench80.5%
MMLU-Pro86.8%
IOI68.4%
LegalBench84.8%
TaxEval72.4%
Terminal-Bench 2.1 (Vals)71.5%
SWE-bench (Vals)95.4%
GPQA Diamond (Vals)88.1%
Vals Index57.0

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

GLM 5.3 vs GLM 5.3 Flash: questions

Is GLM 5.3 better than GLM 5.3 Flash?
GLM 5.3 (max) and GLM 5.3 Flash are level on quality (60.0 vs 60.1). The BenchLeader Index combines every independent quality benchmark; GLM 5.3 Flash is ahead overall as of 2026-09-13, but check the category scores for your use.
Is GLM 5.3 better than GLM 5.3 Flash for coding?
GLM 5.3 scores higher in coding (64 vs 64 on the category index, where 50 is average).
Is GLM 5.3 better than GLM 5.3 Flash for agentic tasks?
GLM 5.3 scores higher in agentic tasks (53 vs 52 on the category index, where 50 is average).
Which is cheaper, GLM 5.3 or GLM 5.3 Flash?
GLM 5.3 Flash is cheaper: $0.119 against $2.15 per million tokens, blended at three input tokens per output token.
Which is faster, GLM 5.3 or GLM 5.3 Flash?
GLM 5.3 Flash streams faster: 107 against 66 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.