BenchLeader

Claude Haiku 5.5 vs Muse Spark 1.3

Verdict
  • Muse Spark 1.3 (max) leads on quality: 67.7 vs 61.3.
  • Claude Haiku 5.5 (high) is stronger in maths.
  • Muse Spark 1.3 (max) is stronger in coding, composite, knowledge, long context, reasoning, agents & tools, human preference, multimodal.
  • Claude Haiku 5.5 (high) is 10× cheaper ($0.200 vs $2.00 per 1M blended).
MetricClaude Haiku 5.5 (high)Muse Spark 1.3 (max)
BenchLeader Index61.367.7
Coding score62.664.6
Composite score72.484.1
Knowledge score64.372.9
Long context score63.866.8
Maths score66.263.6
Reasoning score72.178.1
Agents & tools score–70.3
Human preference score–69.4
Multimodal score–65.9
Blended price $/M$0.200$2.00
Output speed174 tok/s175 tok/s
Time to first answer22.6 s70.5 s
Context window1M1.0M
FrontierMath Tiers 1–3–74.0%
FrontierMath Tier 4–46.3%
OTIS Mock AIME97.2%–
SciCode–58.8%
ProofBench–58.0%
LMArena Text–1494
LMArena Hard Prompts–1517
LMArena Coding–1538
LMArena WebDev15871657
LMArena Vision–1309
LMArena Agent–4
AA Intelligence Index v4.3.237.848.1
AA-LCR77.3%83.0%
AA-Omniscience5.825
GPQA Diamond (AA)–93.5%
Humanity's Last Exam (AA)37.3%48.7%
SciCode (AA)48.7%58.8%
IOI–56.6%
Terminal-Bench 2.1 (Vals)–79.0%
Vals Index–58.2
CritPt18.6%24.9%
GDPval-AA v2.145.9%59.0%
τ³-Banking (AA)–50.5%
ITBench SRE (AA)–33.2%
Code Migration–47.4%
CUA-bench–5.8%
CyberBench–72.7%
Excel Modeling Benchmark–67.4%
Finance Agent v2–60.0%
Harvey's Legal Agent Benchmark–23.8%
Legal Research Bench–55.3%
MysteryMechanism–36.0%
Tax Agent Bench–72.4%
Terminal-Bench 4.0 (Vals)–24.8%
Terminal-Bench Science–10.0%
Vibe Code Bench 1-100–20.5%
Vibe Code Bench v1.1–85.9%
LMArena Maths–1502
LMArena Creative Writing–1456
LMArena Instruction Following–1483
LMArena Multi-turn–1487
LMArena Longer Queries–1497
LMArena Document–1468
Chess Puzzles–38.0%
Mystery Game Puzzles–25.0%
CursorBench–41.6%
Terminal-Bench 4.0 (AA)21.7%33.3%
Terminal-Bench 2.1 (AA)–84.3%
AutomationBench33.7%57.9%
GDP.pdf17.2%26.6%
MLCR–43.3%
Harvey LAB1.1%8.9%
AA-Omniscience: accuracy34.8%43.6%
AA-Omniscience: non-hallucination55.5%67.1%
AA-Briefcase v1.114421581

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Claude Haiku 5.5 vs Muse Spark 1.3: questions

Is Claude Haiku 5.5 better than Muse Spark 1.3?
Muse Spark 1.3 (max) leads on quality: 67.7 vs 61.3. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Claude Haiku 5.5 better than Muse Spark 1.3 for coding?
Muse Spark 1.3 scores higher in coding (65 vs 63 on the category index, where 50 is average).
Which is cheaper, Claude Haiku 5.5 or Muse Spark 1.3?
Claude Haiku 5.5 is cheaper: $0.200 against $2.00 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Haiku 5.5 or Muse Spark 1.3?
Muse Spark 1.3 streams faster: 175 against 174 output tokens per second.
Which has the larger context window?
Muse Spark 1.3 accepts more context: 1.0M against 1M tokens.