BenchLeader

Deepseek v4 Pro vs Muse Spark 1.1

Verdict
  • Muse Spark 1.1 leads on quality: 65.0 vs 60.4.
  • Deepseek v4 Pro (high) is stronger in maths.
  • Muse Spark 1.1 is stronger in coding, human preference, reasoning, agents & tools, composite, instruction following, knowledge, long context, multimodal.
MetricDeepseek v4 Pro (high)Muse Spark 1.1
BenchLeader Index60.465.0
Coding score59.468.1
Human preference score65.866.3
Maths score66.252.2
Reasoning score66.069.2
Agents & tools score62.7
Composite score72.3
Instruction following score77.8
Knowledge score73.2
Long context score65.4
Multimodal score65.0
Blended price $/M$2.00
Output speed194 tok/s
Time to first answer2.6 s
Context window1.0M
GPQA Diamond90.9%
OTIS Mock AIME95.6%
SimpleQA Verified57.8%
SciCode46.4%58.2%
WeirdML46.5%
APEX-Agents41.9%
ProofBench39.0%
Epoch Capabilities Index154.6
LMArena Text14601492
LMArena Hard Prompts14821511
LMArena Coding15051531
LMArena WebDev15801541
LMArena Vision1293
LMArena Agent-2.6
AA Intelligence Index34.3
AA-LCR77.7%
AA-Omniscience28.1
GPQA Diamond (AA)89.8%
Humanity's Last Exam (AA)46.2%
SciCode (AA)58.8%
SWE-Bench Pro61.5%
MCP Atlas88.1%
MultiChallenge75.3%
PRBench Finance55.0%
PRBench Legal57.0%
MultiNRC65.6%
EQ-Bench 41260

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Deepseek v4 Pro vs Muse Spark 1.1: questions

Is Deepseek v4 Pro better than Muse Spark 1.1?
Muse Spark 1.1 leads on quality: 65.0 vs 60.4. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.1 is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Deepseek v4 Pro better than Muse Spark 1.1 for coding?
Muse Spark 1.1 scores higher in coding (68 vs 59 on the category index, where 50 is average).