BenchLeader

Claude Sonnet 4.5 vs Muse Spark

Verdict
  • Muse Spark leads on quality: 65.2 vs 57.9.
  • Muse Spark is stronger in coding, human preference, reasoning, agents & tools, instruction following, knowledge, maths, multimodal.
MetricClaude Sonnet 4.5 (high)Muse Spark
BenchLeader Index57.965.2
Coding score56.466.2
Human preference score65.369.2
Reasoning score66.470.2
Agents & tools score62.9
Instruction following score83.8
Knowledge score63.8
Maths score56.4
Multimodal score66.7
Blended price $/M$6.00
Output speed40 tok/s
Time to first answer1.2 s
Context window1M
GPQA Diamond89.8%
OTIS Mock AIME88.9%
Humanity's Last Exam40.6%
SciCode51.5%
ProofBench17.0%
Epoch Capabilities Index152.1
LMArena Text14561488
LMArena Hard Prompts14871505
LMArena Coding15201526
LMArena WebDev1393
LMArena Vision1306
AIME (Vals)96.9%
MMLU-Pro87.3%
LegalBench84.2%
CorpFin65.1%
TaxEval77.7%
SWE-bench (Vals)74.4%
GPQA Diamond (Vals)89.7%
SWE-Bench Pro55.0%
MCP Atlas82.2%
MultiChallenge75.5%
PRBench Finance52.4%
PRBench Legal52.3%
MultiNRC59.0%
TutorBench68.5%
SWE-bench Verified (bash only)71.4%

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

Claude Sonnet 4.5 vs Muse Spark: questions

Is Claude Sonnet 4.5 better than Muse Spark?
Muse Spark leads on quality: 65.2 vs 57.9. The BenchLeader Index combines every independent quality benchmark; Muse Spark is ahead overall as of 2026-09-13, but check the category scores for your use.
Is Claude Sonnet 4.5 better than Muse Spark for coding?
Muse Spark scores higher in coding (66 vs 56 on the category index, where 50 is average).