BenchLeader

GLM 5.1 vs Muse Spark 1.3

Verdict
  • Muse Spark 1.3 leads on quality: 68.3 vs 60.3.
  • GLM 5.1 is stronger in agents & tools, coding, human preference, instruction following, maths, reasoning.
  • Muse Spark 1.3 is stronger in composite, knowledge, long context.
  • They cost about the same ($2.00 per 1M blended).
  • Muse Spark 1.3 streams 3.1× faster (241 vs 78 tokens per second).
MetricGLM 5.1Muse Spark 1.3
BenchLeader Index60.368.3
Agents & tools score56.9
Coding score57.2
Composite score62.790.3
Human preference score66.4
Instruction following score74.4
Knowledge score56.779.1
Long context score63.468.2
Maths score53.0
Reasoning score61.4
Blended price $/M$2.15$2.00
Output speed78 tok/s241 tok/s
Time to first answer50.4 s30.5 s
Context window200k1.0M
GPQA Diamond89.9%
FrontierMath Tiers 1–336.8%
OTIS Mock AIME93.3%
SWE-bench Verified (Epoch)74.2%
SimpleQA Verified34.0%
SimpleBench55.1%
SciCode43.8%
WeirdML57.1%
ProofBench22.2%
Epoch Capabilities Index149.7
LMArena Text1466
LMArena Hard Prompts1489
LMArena Coding1513
LMArena WebDev1508
AA Intelligence Index26.448.2
IFBench76.3%
AA-LCR73.7%83.0%
AA-Omniscience0.825
Terminal-Bench Hard43.2%
GPQA Diamond (AA)86.8%93.5%
Humanity's Last Exam (AA)30.1%48.7%
SciCode (AA)44.8%58.8%
τ²-Bench Telecom (AA)97.7%
AIME (Vals)91.9%
LiveCodeBench81.4%
MMLU-Pro86.9%
LegalBench84.4%
CorpFin64.5%
TaxEval71.2%
Terminal-Bench 2.1 (Vals)56.9%
SWE-bench (Vals)76.4%
GPQA Diamond (Vals)84.5%
AIME 202695.8%
HMMT February 202689.4%
MathArena Apex11.5%
MCP Atlas75.6%
PRBench Finance59.5%
PRBench Legal61.6%

Data as of 2026-09-14. Best configuration of each model; every score links to its source on the model pages.

GLM 5.1 vs Muse Spark 1.3: questions

Is GLM 5.1 better than Muse Spark 1.3?
Muse Spark 1.3 leads on quality: 68.3 vs 60.3. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 is ahead overall as of 2026-09-14, but check the category scores for your use.
Which is cheaper, GLM 5.1 or Muse Spark 1.3?
Muse Spark 1.3 is cheaper: $2.00 against $2.15 per million tokens, blended at three input tokens per output token.
Which is faster, GLM 5.1 or Muse Spark 1.3?
Muse Spark 1.3 streams faster: 241 against 78 output tokens per second.
Which has the larger context window?
Muse Spark 1.3 accepts more context: 1.0M against 200k tokens.