BenchLeader

DeepSeek V4 Flash vs Muse Spark

Verdict
  • Muse Spark leads on quality: 65.6 vs 59.1.
  • Muse Spark is stronger in agents & tools, coding, composite, human preference, instruction following, knowledge, long context, reasoning, maths, multimodal.
MetricDeepSeek V4 Flash (high)Muse Spark
BenchLeader Index59.165.6
Agents & tools score60.066.4
Coding score57.366.2
Composite score60.468.6
Human preference score63.269.3
Instruction following score71.579.7
Knowledge score53.264.6
Long context score62.465.5
Reasoning score63.270.2
Maths score56.4
Multimodal score65.9
Blended price $/M$0.168
Output speed44 tok/s
Time to first answer11.7 s
Context window1M262k
GPQA Diamond89.8%
OTIS Mock AIME88.9%
Humanity's Last Exam40.6%
SciCode42.0%51.5%
WeirdML43.8%
ProofBench17.0%
Epoch Capabilities Index152.1
LMArena Text14391488
LMArena Hard Prompts14581505
LMArena Coding14791526
LMArena WebDev1582
LMArena Vision1306
LMArena Agent2.2
AA Intelligence Index24.831.3
IFBench73.5%75.9%
AA-LCR72.0%78.0%
MMMU-Pro80.5%
AA-Omniscience-23.17.2
Terminal-Bench Hard38.6%45.5%
GPQA Diamond (AA)86.7%88.4%
Humanity's Last Exam (AA)30.3%40.7%
SciCode (AA)40.2%
τ²-Bench Telecom (AA)95.6%91.5%
AIME (Vals)96.9%
MMLU-Pro87.3%
LegalBench84.2%
CorpFin65.1%
TaxEval77.7%
SWE-bench (Vals)74.4%
GPQA Diamond (Vals)89.7%
SWE-Bench Pro55.0%
MCP Atlas82.2%
MultiChallenge75.5%
PRBench Finance52.4%
PRBench Legal52.3%
MultiNRC59.0%
TutorBench68.5%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

DeepSeek V4 Flash vs Muse Spark: questions

Is DeepSeek V4 Flash better than Muse Spark?
Muse Spark leads on quality: 65.6 vs 59.1. The BenchLeader Index combines every independent quality benchmark; Muse Spark is ahead overall as of 2026-09-10, but check the category scores for your use.
Is DeepSeek V4 Flash better than Muse Spark for coding?
Muse Spark scores higher in coding (66 vs 57 on the category index, where 50 is average).
Is DeepSeek V4 Flash better than Muse Spark for agentic tasks?
Muse Spark scores higher in agentic tasks (66 vs 60 on the category index, where 50 is average).
Which has the larger context window?
DeepSeek V4 Flash accepts more context: 1M against 262k tokens.