BenchLeader

Claude Sonnet 4.6 vs Muse Spark 1.1

Verdict
  • Muse Spark 1.1 leads on quality: 65.0 vs 60.4.
  • Claude Sonnet 4.6 is stronger in long context, maths.
  • Muse Spark 1.1 is stronger in agents & tools, coding, composite, human preference, instruction following, knowledge, multimodal, reasoning.
  • Muse Spark 1.1 is 3.0× cheaper ($2.00 vs $6.00 per 1M blended).
  • Muse Spark 1.1 streams 4.6× faster (194 vs 42 tokens per second).
MetricClaude Sonnet 4.6Muse Spark 1.1
BenchLeader Index60.465.0
Agents & tools score54.662.7
Coding score57.168.1
Composite score67.572.3
Human preference score62.066.3
Instruction following score56.977.8
Knowledge score62.373.2
Long context score66.665.4
Maths score62.752.2
Multimodal score60.665.0
Reasoning score65.269.2
Blended price $/M$6.00$2.00
Output speed42 tok/s194 tok/s
Time to first answer2.2 s2.6 s
Context window1M1.0M
GPQA Diamond87.4%
OTIS Mock AIME85.8%
SWE-bench Verified (Epoch)75.2%
SimpleQA Verified57.8%
Terminal-Bench53.4%
SciCode58.2%
APEX-Agents41.9%
FrontierCode24.3%
ProofBench39.0%
Epoch Capabilities Index152.3154.6
LMArena Text14721492
LMArena Hard Prompts15041511
LMArena Coding15281531
LMArena WebDev15211541
LMArena Vision12821293
LMArena Agent-1-2.6
AA Intelligence Index30.434.3
IFBench56.6%
AA-LCR80.0%77.7%
MMMU-Pro73.3%
AA-Omniscience12.228.1
Terminal-Bench Hard53.0%
GPQA Diamond (AA)87.5%89.8%
Humanity's Last Exam (AA)33.6%46.2%
SciCode (AA)50.1%58.8%
τ²-Bench Telecom (AA)79.5%
AIME (Vals)92.3%
LiveCodeBench82.1%
MMLU-Pro87.3%
LegalBench82.1%
CorpFin65.3%
TaxEval77.1%
MedQA92.1%
Terminal-Bench 2.1 (Vals)57.3%
SWE-bench (Vals)77.4%
GPQA Diamond (Vals)85.6%
Vals Index50.6
SWE-Bench Pro61.5%
MCP Atlas69.5%88.1%
MultiChallenge75.3%
PRBench Finance55.0%
PRBench Legal57.0%
MultiNRC65.6%
EQ-Bench 412071260

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Claude Sonnet 4.6 vs Muse Spark 1.1: questions

Is Claude Sonnet 4.6 better than Muse Spark 1.1?
Muse Spark 1.1 leads on quality: 65.0 vs 60.4. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.1 is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Claude Sonnet 4.6 better than Muse Spark 1.1 for coding?
Muse Spark 1.1 scores higher in coding (68 vs 57 on the category index, where 50 is average).
Is Claude Sonnet 4.6 better than Muse Spark 1.1 for agentic tasks?
Muse Spark 1.1 scores higher in agentic tasks (63 vs 55 on the category index, where 50 is average).
Which is cheaper, Claude Sonnet 4.6 or Muse Spark 1.1?
Muse Spark 1.1 is cheaper: $2.00 against $6.00 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Sonnet 4.6 or Muse Spark 1.1?
Muse Spark 1.1 streams faster: 194 against 42 output tokens per second.
Which has the larger context window?
Muse Spark 1.1 accepts more context: 1.0M against 1M tokens.