BenchLeader

GLM 5.2 vs Muse Spark

Verdict
  • Muse Spark leads on quality: 65.2 vs 57.7.
  • Muse Spark is stronger in agents & tools, coding, human preference, knowledge, maths, reasoning, instruction following, multimodal.
MetricGLM 5.2 (max)Muse Spark
BenchLeader Index57.765.2
Agents & tools score56.962.9
Coding score63.866.2
Human preference score67.369.2
Knowledge score44.863.8
Maths score55.856.4
Reasoning score66.870.2
Instruction following score83.8
Multimodal score66.7
Blended price $/M$2.15
Output speed75 tok/s
Time to first answer34.2 s
Context window1M
GPQA Diamond91.9%89.8%
FrontierMath Tiers 1–359.2%
FrontierMath Tier 429.3%
OTIS Mock AIME86.4%88.9%
SWE-bench Verified (Epoch)78.7%
SimpleQA Verified34.2%
Humanity's Last Exam40.6%
SciCode50.5%51.5%
WeirdML70.1%
ProofBench35.0%17.0%
Epoch Capabilities Index152.1
LMArena Text14721488
LMArena Hard Prompts14931505
LMArena Coding15101526
LMArena WebDev1592
LMArena Vision1306
LMArena Agent4.6
AIME (Vals)96.9%
MMLU-Pro87.3%
LegalBench84.2%
CorpFin65.1%
TaxEval77.7%
Terminal-Bench 2.1 (Vals)67.8%
SWE-bench (Vals)82.8%74.4%
GPQA Diamond (Vals)89.7%
SWE-Bench Pro55.0%
MCP Atlas82.2%
MultiChallenge75.5%
PRBench Finance52.4%
PRBench Legal52.3%
MultiNRC59.0%
TutorBench68.5%

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

GLM 5.2 vs Muse Spark: questions

Is GLM 5.2 better than Muse Spark?
Muse Spark leads on quality: 65.2 vs 57.7. The BenchLeader Index combines every independent quality benchmark; Muse Spark is ahead overall as of 2026-09-13, but check the category scores for your use.
Is GLM 5.2 better than Muse Spark for coding?
Muse Spark scores higher in coding (66 vs 64 on the category index, where 50 is average).
Is GLM 5.2 better than Muse Spark for agentic tasks?
Muse Spark scores higher in agentic tasks (63 vs 57 on the category index, where 50 is average).