BenchLeader

Claude Opus 5.5 vs GLM-5

Verdict
  • Claude Opus 5.5 (thinking) leads on quality: 70.7 vs 59.4.
  • Claude Opus 5.5 (thinking) is stronger in composite, knowledge, long context, multimodal, reasoning.
  • GLM-5 (thinking) is stronger in agents & tools, coding, instruction following, maths.
  • GLM-5 (thinking) is 5.2× cheaper ($1.55 vs $8.00 per 1M blended).
  • GLM-5 (thinking) streams 1.5× faster (70 vs 48 tokens per second).
MetricClaude Opus 5.5 (thinking)GLM-5 (thinking)
BenchLeader Index70.759.4
Composite score95.063.2
Knowledge score85.058.9
Long context score68.563.9
Multimodal score71.9
Reasoning score95.053.1
Agents & tools score57.9
Coding score59.2
Instruction following score70.7
Maths score63.2
Blended price $/M$8.00$1.55
Output speed48 tok/s70 tok/s
Time to first answer4.2 s46.1 s
Context window1M205k
AA Intelligence Index57.627.9
IFBench72.3%
AA-LCR84.7%75.7%
MMMU-Pro87.7%
AA-Omniscience46.40.3
Terminal-Bench Hard43.2%
GPQA Diamond (AA)82.0%
Humanity's Last Exam (AA)61.4%29.3%
SciCode (AA)66.9%
τ²-Bench Telecom (AA)98.3%
AIME (Vals)91.7%
LiveCodeBench81.9%
MMLU-Pro86.0%
LegalBench84.1%
CorpFin62.9%
TaxEval70.0%
MedQA94.3%
SWE-bench (Vals)71.4%
GPQA Diamond (Vals)83.3%
Kagi LLM Benchmark75.0%
τ²-bench81.0%
CritPt31.7%2.0%
GDPval (AA)67.3%
APEX-Agents (AA)14.4%
CaseLaw v252.5%
Terminal-Bench 2.0 (Vals)49.4%
Vibe Code Bench v1.123.4%
Terminal-Bench 4.0 (AA)59.6%
AutomationBench69.5%
GDP.pdf26.2%
Harvey LAB91.2%
AA-Omniscience: accuracy66.2%26.3%
AA-Omniscience: non-hallucination41.4%64.7%
AA-Briefcase1822

Data as of 2026-09-22. Best configuration of each model; every score links to its source on the model pages.

Claude Opus 5.5 vs GLM-5: questions

Is Claude Opus 5.5 better than GLM-5?
Claude Opus 5.5 (thinking) leads on quality: 70.7 vs 59.4. The BenchLeader Index combines every independent quality benchmark; Claude Opus 5.5 (thinking) is ahead overall as of 2026-09-22, but check the category scores for your use.
Which is cheaper, Claude Opus 5.5 or GLM-5?
GLM-5 is cheaper: $1.55 against $8.00 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Opus 5.5 or GLM-5?
GLM-5 streams faster: 70 against 48 output tokens per second.
Which has the larger context window?
Claude Opus 5.5 accepts more context: 1M against 205k tokens.