BenchLeader

Claude Opus 4.8 vs Muse Spark

Verdict
  • Claude Opus 4.8 and Muse Spark are level on quality (64.7 vs 65.4).
  • Claude Opus 4.8 is stronger in composite, knowledge.
  • Muse Spark is stronger in agents & tools, coding, human preference, instruction following, long context, multimodal, reasoning, maths.
MetricClaude Opus 4.8Muse Spark
BenchLeader Index64.765.4
Agents & tools score65.066.3
Coding score65.766.0
Composite score81.968.4
Human preference score65.669.0
Instruction following score61.679.6
Knowledge score66.364.5
Long context score65.365.5
Multimodal score64.465.9
Reasoning score64.270.0
Maths score56.3
Blended price $/M$10.00
Output speed58 tok/s
Time to first answer29.5 s
Context window1M262k
GPQA Diamond89.8%
OTIS Mock AIME88.9%
Humanity's Last Exam40.6%
SimpleBench64.8%
SciCode51.5%
Remote Labor Index8.3%
FrontierCode46.5%
ProofBench17.0%
GSO-Bench47.1%
Epoch Capabilities Index158.3152.1
LMArena Text14731488
LMArena Hard Prompts15031505
LMArena Coding15271526
LMArena WebDev1540
LMArena Vision12891306
AA Intelligence Index42.031.3
IFBench62.2%75.9%
AA-LCR77.7%78.0%
MMMU-Pro80.5%
AA-Omniscience28.87.2
Terminal-Bench Hard58.3%45.5%
GPQA Diamond (AA)92.0%88.4%
Humanity's Last Exam (AA)48.7%40.7%
SciCode (AA)54.4%
τ²-Bench Telecom (AA)94.4%91.5%
AIME (Vals)96.9%
LiveCodeBench87.8%
MMLU-Pro89.6%87.3%
LegalBench83.6%84.2%
CorpFin66.7%65.1%
TaxEval75.6%77.7%
Terminal-Bench 2.1 (Vals)71.9%
SWE-bench (Vals)88.6%74.4%
GPQA Diamond (Vals)92.4%89.7%
Vals Index60.9
SWE-Bench Pro55.0%
MCP Atlas82.2%
MultiChallenge75.5%
PRBench Finance52.4%
PRBench Legal52.3%
MultiNRC59.0%
HiL-Bench35.3%
TutorBench68.5%
EQ-Bench 41281

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.