BenchLeader

GLM 5.1 vs Muse Spark

Verdict
  • Muse Spark leads on quality: 65.6 vs 60.2.
  • Muse Spark is stronger in agents & tools, coding, composite, human preference, instruction following, knowledge, long context, maths, reasoning, multimodal.
MetricGLM 5.1Muse Spark
BenchLeader Index60.265.6
Agents & tools score56.866.4
Coding score57.366.2
Composite score62.468.6
Human preference score66.669.3
Instruction following score73.979.7
Knowledge score56.664.6
Long context score63.365.5
Maths score53.056.4
Reasoning score61.470.2
Multimodal score65.9
Blended price $/M$2.15
Output speed65 tok/s
Time to first answer60.1 s
Context window200k262k
GPQA Diamond89.9%89.8%
FrontierMath Tiers 1–336.8%
OTIS Mock AIME93.3%88.9%
SWE-bench Verified (Epoch)74.2%
SimpleQA Verified34.0%
Humanity's Last Exam40.6%
SimpleBench55.1%
SciCode43.8%51.5%
WeirdML57.1%
ProofBench22.2%17.0%
Epoch Capabilities Index149.7152.1
LMArena Text14661488
LMArena Hard Prompts14891505
LMArena Coding15141526
LMArena WebDev1508
LMArena Vision1306
AA Intelligence Index26.431.3
IFBench76.3%75.9%
AA-LCR73.7%78.0%
MMMU-Pro80.5%
AA-Omniscience0.87.2
Terminal-Bench Hard43.2%45.5%
GPQA Diamond (AA)86.8%88.4%
Humanity's Last Exam (AA)30.1%40.7%
SciCode (AA)44.8%
τ²-Bench Telecom (AA)97.7%91.5%
AIME (Vals)91.9%96.9%
LiveCodeBench81.4%
MMLU-Pro86.9%87.3%
LegalBench84.4%84.2%
CorpFin64.5%65.1%
TaxEval71.2%77.7%
Terminal-Bench 2.1 (Vals)56.9%
SWE-bench (Vals)76.4%74.4%
GPQA Diamond (Vals)84.5%89.7%
AIME 202695.8%
HMMT February 202689.4%
MathArena Apex11.5%
SWE-Bench Pro55.0%
MCP Atlas75.6%82.2%
MultiChallenge75.5%
PRBench Finance52.4%
PRBench Legal52.3%
MultiNRC59.0%
TutorBench68.5%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

GLM 5.1 vs Muse Spark: questions

Is GLM 5.1 better than Muse Spark?
Muse Spark leads on quality: 65.6 vs 60.2. The BenchLeader Index combines every independent quality benchmark; Muse Spark is ahead overall as of 2026-09-10, but check the category scores for your use.
Is GLM 5.1 better than Muse Spark for coding?
Muse Spark scores higher in coding (66 vs 57 on the category index, where 50 is average).
Is GLM 5.1 better than Muse Spark for agentic tasks?
Muse Spark scores higher in agentic tasks (66 vs 57 on the category index, where 50 is average).
Which has the larger context window?
Muse Spark accepts more context: 262k against 200k tokens.