BenchLeader

GLM-5.2 vs Muse Spark 1.3

Verdict
  • Muse Spark 1.3 (max) leads on quality: 69.3 vs 63.9.
  • GLM-5.2 (max) is stronger in instruction following.
  • Muse Spark 1.3 (max) is stronger in agents & tools, coding, composite, human preference, knowledge, long context, maths, reasoning, multimodal.
  • They cost about the same ($2.00 per 1M blended).
  • Muse Spark 1.3 (max) streams 3.1× faster (223 vs 71 tokens per second).
MetricGLM-5.2 (max)Muse Spark 1.3 (max)
BenchLeader Index63.969.3
Agents & tools score63.871.1
Coding score63.866.6
Composite score72.190.1
Human preference score67.269.8
Instruction following score71.6
Knowledge score55.576.1
Long context score65.668.0
Maths score59.069.2
Reasoning score72.983.1
Multimodal score67.4
Blended price $/M$2.15$2.00
Output speed71 tok/s223 tok/s
Time to first answer31.6 s34.8 s
Context window1M1.0M
GPQA Diamond91.9%
FrontierMath Tiers 1–359.2%
FrontierMath Tier 429.3%
OTIS Mock AIME86.4%
SWE-bench Verified (Epoch)78.7%
SimpleQA Verified34.2%
SciCode50.5%
WeirdML70.1%
ProofBench35.0%
LMArena Text14721493
LMArena Hard Prompts14931516
LMArena Coding15101537
LMArena WebDev15921652
LMArena Vision1315
LMArena Agent4.44.2
AA Intelligence Index34.048.2
IFBench73.3%
AA-LCR78.3%83.0%
AA-Omniscience4.425
Terminal-Bench Hard50.8%
GPQA Diamond (AA)89.5%93.5%
Humanity's Last Exam (AA)41.1%48.7%
SciCode (AA)51.2%58.8%
τ²-Bench Telecom (AA)99.1%
IOI56.6%
Terminal-Bench 2.1 (Vals)67.8%79.0%
SWE-bench (Vals)82.8%
Vals Index64.5
CritPt20.9%24.9%
GDPval (AA)45.3%60.2%
τ²-Bench Banking (AA)34.6%50.5%
ITBench SRE (AA)42.7%
APEX-Agents (AA)33.7%
Code Migration37.9%47.4%
Excel Modeling Benchmark67.4%
Finance Agent v260.0%
Harvey's Legal Agent Benchmark7.1%23.8%
Legal Research Bench31.3%55.3%
MysteryMechanism36.0%
ProgramBench0.5%
SkillsBench45.1%
SREBench0.0%
Terminal-Bench 4.0 (Vals)27.8%
Terminal-Bench Science14.3%
Vibe Code Bench 1-10020.5%
Vibe Code Bench v1.164.0%85.9%
LMArena Maths14801496
LMArena Creative Writing14511453
LMArena Instruction Following14661483
LMArena Multi-turn14691493
LMArena Longer Queries14831506
LMArena Document1468
Chess Puzzles21.0%
EBR-bench9.5%
PostTrainBench31.7%
DeepSWE43.8%
LMCA45.8%
DTBench93.6%
CursorBench55.0%
ALE-Bench1010.2

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

GLM-5.2 vs Muse Spark 1.3: questions

Is GLM-5.2 better than Muse Spark 1.3?
Muse Spark 1.3 (max) leads on quality: 69.3 vs 63.9. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 (max) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is GLM-5.2 better than Muse Spark 1.3 for coding?
Muse Spark 1.3 scores higher in coding (67 vs 64 on the category index, where 50 is average).
Is GLM-5.2 better than Muse Spark 1.3 for agentic tasks?
Muse Spark 1.3 scores higher in agentic tasks (71 vs 64 on the category index, where 50 is average).
Which is cheaper, GLM-5.2 or Muse Spark 1.3?
Muse Spark 1.3 is cheaper: $2.00 against $2.15 per million tokens, blended at three input tokens per output token.
Which is faster, GLM-5.2 or Muse Spark 1.3?
Muse Spark 1.3 streams faster: 223 against 71 output tokens per second.
Which has the larger context window?
Muse Spark 1.3 accepts more context: 1.0M against 1M tokens.