BenchLeader

GLM 5.3 vs Muse Spark 1.1

Verdict
  • Muse Spark 1.1 leads on quality: 65.0 vs 60.0.
  • GLM 5.3 (max) is stronger in human preference, maths.
  • Muse Spark 1.1 is stronger in agents & tools, coding, knowledge, reasoning, composite, instruction following, long context, multimodal.
  • They cost about the same ($2.00 per 1M blended).
  • Muse Spark 1.1 streams 3.5× faster (194 vs 56 tokens per second).
MetricGLM 5.3 (max)Muse Spark 1.1
BenchLeader Index60.065.0
Agents & tools score52.962.7
Coding score65.268.1
Human preference score68.666.3
Knowledge score55.873.2
Maths score60.052.2
Reasoning score66.969.2
Composite score72.3
Instruction following score77.8
Long context score65.4
Multimodal score65.0
Blended price $/M$2.15$2.00
Output speed56 tok/s194 tok/s
Time to first answer37.4 s2.6 s
Context window1.3M1.0M
GPQA Diamond90.9%
FrontierMath Tiers 1–368.8%
FrontierMath Tier 429.3%
OTIS Mock AIME91.1%
SimpleQA Verified41.0%57.8%
Terminal-Bench41.8%
SciCode56.5%58.2%
WeirdML75.4%
APEX-Agents41.9%
ProofBench49.0%39.0%
Epoch Capabilities Index154.6
LMArena Text14821492
LMArena Hard Prompts15071511
LMArena Coding15281531
LMArena WebDev16131541
LMArena Vision1293
LMArena Agent2.6-2.6
AA Intelligence Index34.3
AA-LCR77.7%
AA-Omniscience28.1
GPQA Diamond (AA)89.8%
Humanity's Last Exam (AA)46.2%
SciCode (AA)58.8%
LiveCodeBench80.5%
MMLU-Pro86.8%
LegalBench84.8%
TaxEval72.4%
Terminal-Bench 2.1 (Vals)71.5%
SWE-bench (Vals)95.4%
GPQA Diamond (Vals)88.1%
Vals Index57.0
SWE-Bench Pro61.5%
MCP Atlas88.1%
MultiChallenge75.3%
PRBench Finance55.0%
PRBench Legal57.0%
MultiNRC65.6%
EQ-Bench 41260

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

GLM 5.3 vs Muse Spark 1.1: questions

Is GLM 5.3 better than Muse Spark 1.1?
Muse Spark 1.1 leads on quality: 65.0 vs 60.0. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.1 is ahead overall as of 2026-09-10, but check the category scores for your use.
Is GLM 5.3 better than Muse Spark 1.1 for coding?
Muse Spark 1.1 scores higher in coding (68 vs 65 on the category index, where 50 is average).
Is GLM 5.3 better than Muse Spark 1.1 for agentic tasks?
Muse Spark 1.1 scores higher in agentic tasks (63 vs 53 on the category index, where 50 is average).
Which is cheaper, GLM 5.3 or Muse Spark 1.1?
Muse Spark 1.1 is cheaper: $2.00 against $2.15 per million tokens, blended at three input tokens per output token.
Which is faster, GLM 5.3 or Muse Spark 1.1?
Muse Spark 1.1 streams faster: 194 against 56 output tokens per second.
Which has the larger context window?
GLM 5.3 accepts more context: 1.3M against 1.0M tokens.