BenchLeader

GLM 5.2 vs Qwen3.8 Max

Verdict
  • Qwen3.8 Max (max) leads on quality: 64.2 vs 62.6.
  • GLM 5.2 (max) is stronger in agents & tools, instruction following.
  • Qwen3.8 Max (max) is stronger in coding, composite, human preference, knowledge, long context, maths, reasoning, multimodal.
  • GLM 5.2 (max) is 1.4× cheaper ($2.15 vs $3.00 per 1M blended).
  • GLM 5.2 (max) streams 2.2× faster (82 vs 37 tokens per second).
MetricGLM 5.2 (max)Qwen3.8 Max (max)
BenchLeader Index62.664.2
Agents & tools score62.958.2
Coding score62.464.6
Composite score67.770.7
Human preference score67.068.0
Instruction following score71.6–
Knowledge score54.062.7
Long context score64.365.4
Maths score57.765.1
Reasoning score70.272.2
Multimodal score–66.2
Blended price $/M$2.15$3.00
Output speed82 tok/s37 tok/s
Time to first answer28.4 s58.9 s
Context window1M1M
GPQA Diamond91.9%–
FrontierMath Tiers 1–359.2%–
FrontierMath Tier 429.3%–
OTIS Mock AIME86.4%–
SWE-bench Verified (Epoch)78.7%–
SimpleQA Verified34.2%–
Terminal-Bench–27.0%
SciCode50.5%53.2%
WeirdML70.1%–
APEX-Agents–63.3%
ProofBench35.0%58.0%
Epoch Capabilities Index–156.4
LMArena Text14751483
LMArena Hard Prompts14951504
LMArena Coding15121524
LMArena WebDev16031672
LMArena Vision–1314
LMArena Agent3.22.3
LiveBench–78.5%
LiveBench Reasoning–88.2%
LiveBench Coding–72.9%
LiveBench Agentic Coding–64.7%
LiveBench Mathematics–91.3%
LiveBench Data Analysis–78.4%
LiveBench Language–79.7%
LiveBench Instruction Following–74.1%
AA Intelligence Index v4.3.233.745.4
IFBench73.3%–
AA-LCR78.3%80.3%
MMMU-Pro–82.8%
AA-Omniscience4.412.0
Terminal-Bench Hard50.8%–
GPQA Diamond (AA)89.5%92.8%
Humanity's Last Exam (AA)41.1%43.1%
SciCode (AA)51.2%53.2%
τ²-Bench Telecom (AA)99.1%–
LiveCodeBench–87.8%
MMLU-Pro–88.6%
IOI–68.9%
LegalBench–83.6%
CorpFin–65.8%
TaxEval–75.5%
Terminal-Bench 2.1 (Vals)67.8%67.4%
SWE-bench (Vals)82.8%85.6%
GPQA Diamond (Vals)–93.7%
Vals Index–48.3
CritPt20.9%20.0%
GDPval-AA v2.143.7%58.6%
τ³-Banking (AA)34.6%51.3%
ITBench SRE (AA)42.7%40.3%
Analyst Agent (AA)–45.0%
APEX-Agents (AA)33.7%42.4%
Code Migration37.9%24.0%
CyberBench–28.6%
Excel Modeling Benchmark–60.1%
Finance Agent v2–50.6%
Harvey's Legal Agent Benchmark–10.4%
Legal Research Bench31.3%47.6%
MedCode–40.7%
MedScribe–85.0%
MMMU-Pro (Vals)–88.0%
MortgageTax–64.0%
MysteryMechanism–23.9%
ProgramBench0.5%0.0%
Public Benefits Bench–67.1%
SAGE–51.3%
SkillsBench45.1%42.0%
SREBench0.0%–
Tax Agent Bench–66.0%
Terminal-Bench 4.0 (Vals)–34.3%
Terminal-Bench Science–1.4%
Vals Multimodal Index–65.4%
Vibe Code Bench 1-100–12.8%
Vibe Code Bench v1.164.0%64.7%
LMArena Maths14861497
LMArena Creative Writing14531470
LMArena Instruction Following14691474
LMArena Multi-turn14751492
LMArena Longer Queries14861492
Chess Puzzles21.0%–
EBR-bench9.5%–
PostTrainBench31.7%–
DeepSWE v1.143.8%–
LMCA45.8%–
DTBench93.6%–
ALE-Bench1010.2–
Terminal-Bench 4.0 (AA)1.0%38.9%
Terminal-Bench 2.1 (AA)77.9%88.8%
AutomationBench–56.2%
GDP.pdf–22.8%
MLCR–20.0%
EnterpriseOps-Gym–47.6%
AA-Omniscience: accuracy24.3%31.9%
AA-Omniscience: non-hallucination73.7%71.2%
AA-Briefcase v1.1–1617

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

GLM 5.2 vs Qwen3.8 Max: questions

Is GLM 5.2 better than Qwen3.8 Max?
Qwen3.8 Max (max) leads on quality: 64.2 vs 62.6. The BenchLeader Index combines every independent quality benchmark; Qwen3.8 Max (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is GLM 5.2 better than Qwen3.8 Max for coding?
Qwen3.8 Max scores higher in coding (65 vs 62 on the category index, where 50 is average).
Is GLM 5.2 better than Qwen3.8 Max for agentic tasks?
GLM 5.2 scores higher in agentic tasks (63 vs 58 on the category index, where 50 is average).
Which is cheaper, GLM 5.2 or Qwen3.8 Max?
GLM 5.2 is cheaper: $2.15 against $3.00 per million tokens, blended at three input tokens per output token.
Which is faster, GLM 5.2 or Qwen3.8 Max?
GLM 5.2 streams faster: 82 against 37 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.