BenchLeader

Claude Sonnet 4.6 vs Muse Spark

Verdict
  • Muse Spark leads on quality: 65.6 vs 60.4.
  • Claude Sonnet 4.6 is stronger in long context, maths.
  • Muse Spark is stronger in agents & tools, coding, composite, human preference, instruction following, knowledge, multimodal, reasoning.
MetricClaude Sonnet 4.6Muse Spark
BenchLeader Index60.465.6
Agents & tools score54.666.4
Coding score57.166.2
Composite score67.568.6
Human preference score62.069.3
Instruction following score56.979.7
Knowledge score62.364.6
Long context score66.665.5
Maths score62.756.4
Multimodal score60.665.9
Reasoning score65.270.2
Blended price $/M$6.00
Output speed42 tok/s
Time to first answer2.2 s
Context window1M262k
GPQA Diamond87.4%89.8%
OTIS Mock AIME85.8%88.9%
SWE-bench Verified (Epoch)75.2%
Humanity's Last Exam40.6%
Terminal-Bench53.4%
SciCode51.5%
FrontierCode24.3%
ProofBench17.0%
Epoch Capabilities Index152.3152.1
LMArena Text14721488
LMArena Hard Prompts15041505
LMArena Coding15281526
LMArena WebDev1521
LMArena Vision12821306
LMArena Agent-1
AA Intelligence Index30.431.3
IFBench56.6%75.9%
AA-LCR80.0%78.0%
MMMU-Pro73.3%80.5%
AA-Omniscience12.27.2
Terminal-Bench Hard53.0%45.5%
GPQA Diamond (AA)87.5%88.4%
Humanity's Last Exam (AA)33.6%40.7%
SciCode (AA)50.1%
τ²-Bench Telecom (AA)79.5%91.5%
AIME (Vals)92.3%96.9%
LiveCodeBench82.1%
MMLU-Pro87.3%87.3%
LegalBench82.1%84.2%
CorpFin65.3%65.1%
TaxEval77.1%77.7%
MedQA92.1%
Terminal-Bench 2.1 (Vals)57.3%
SWE-bench (Vals)77.4%74.4%
GPQA Diamond (Vals)85.6%89.7%
Vals Index50.6
SWE-Bench Pro55.0%
MCP Atlas69.5%82.2%
MultiChallenge75.5%
PRBench Finance52.4%
PRBench Legal52.3%
MultiNRC59.0%
TutorBench68.5%
EQ-Bench 41207

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Claude Sonnet 4.6 vs Muse Spark: questions

Is Claude Sonnet 4.6 better than Muse Spark?
Muse Spark leads on quality: 65.6 vs 60.4. The BenchLeader Index combines every independent quality benchmark; Muse Spark is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Claude Sonnet 4.6 better than Muse Spark for coding?
Muse Spark scores higher in coding (66 vs 57 on the category index, where 50 is average).
Is Claude Sonnet 4.6 better than Muse Spark for agentic tasks?
Muse Spark scores higher in agentic tasks (66 vs 55 on the category index, where 50 is average).
Which has the larger context window?
Claude Sonnet 4.6 accepts more context: 1M against 262k tokens.