BenchLeader

Deepseek v4 Pro vs Muse Spark 1.3

Verdict
  • Muse Spark 1.3 (xhigh) leads on quality: 62.0 vs 60.4.
  • Deepseek v4 Pro (high) is stronger in human preference, maths, reasoning.
  • Muse Spark 1.3 (xhigh) is stronger in coding, agents & tools, composite, knowledge, long context, multimodal.
MetricDeepseek v4 Pro (high)Muse Spark 1.3 (xhigh)
BenchLeader Index60.462.0
Coding score59.360.5
Human preference score66.0
Maths score66.2
Reasoning score65.9
Agents & tools score60.2
Composite score75.0
Knowledge score71.8
Long context score65.8
Multimodal score65.7
Blended price $/M$2.00
Output speed235 tok/s
Time to first answer28.5 s
Context window1.0M
GPQA Diamond90.9%
OTIS Mock AIME95.6%
SciCode46.4%
WeirdML46.5%
LMArena Text1462
LMArena Hard Prompts1482
LMArena Coding1505
LMArena WebDev15811623
LiveBench81.6%
LiveBench Reasoning89.7%
LiveBench Coding81.1%
LiveBench Agentic Coding64.1%
LiveBench Mathematics96.0%
LiveBench Data Analysis79.6%
LiveBench Language82.8%
AA Intelligence Index45.2
AA-LCR83.0%
MMMU-Pro82.0%
AA-Omniscience23.1
GPQA Diamond (AA)94.1%
Humanity's Last Exam (AA)47.5%
SciCode (AA)59.7%
IOI43.9%
Terminal-Bench 2.1 (Vals)72.3%
Vals Index60.3

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

Deepseek v4 Pro vs Muse Spark 1.3: questions

Is Deepseek v4 Pro better than Muse Spark 1.3?
Muse Spark 1.3 (xhigh) leads on quality: 62.0 vs 60.4. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 (xhigh) is ahead overall as of 2026-09-13, but check the category scores for your use.
Is Deepseek v4 Pro better than Muse Spark 1.3 for coding?
Muse Spark 1.3 scores higher in coding (61 vs 59 on the category index, where 50 is average).