BenchLeader

Claude Fable 5.1 vs Muse Spark 1.2

Verdict
  • Claude Fable 5.1 (xhigh) leads on quality: 70.3 vs 62.6.
  • Claude Fable 5.1 (xhigh) is stronger in coding, composite, knowledge, long context, reasoning.
  • Muse Spark 1.2 is stronger in maths.
  • Muse Spark 1.2 is 10× cheaper ($2.00 vs $20.00 per 1M blended).
  • Muse Spark 1.2 streams 3.8× faster (219 vs 58 tokens per second).
MetricClaude Fable 5.1 (xhigh)Muse Spark 1.2
BenchLeader Index70.362.6
Coding score71.366.5
Composite score95.079.1
Knowledge score84.176.9
Long context score68.166.0
Reasoning score76.470.2
Maths score54.3
Blended price $/M$20.00$2.00
Output speed58 tok/s219 tok/s
Time to first answer132.0 s23.7 s
Context window1M1.0M
Humanity's Last Exam46.5%
SimpleBench74.5%
SciCode60.1%
ProofBench43.0%
Epoch Capabilities Index155.5
AA Intelligence Index53.239.8
AA-LCR83.0%79.0%
AA-Omniscience42.427.2
GPQA Diamond (AA)93.4%90.4%
Humanity's Last Exam (AA)58.7%45.5%
SciCode (AA)60.9%57.4%
IOI49.5%
ARC-AGI-196.5%
ARC-AGI-290.0%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.