BenchLeader

Muse Spark 1.1 vs Qwen3.8-Flash-Next

Verdict
  • Muse Spark 1.1 leads on quality: 65.0 vs 59.1.
  • Muse Spark 1.1 is stronger in agents & tools, composite, human preference, instruction following, knowledge, maths, multimodal, reasoning.
  • Qwen3.8-Flash-Next is stronger in coding, long context.
  • Qwen3.8-Flash-Next is 8.7× cheaper ($0.230 vs $2.00 per 1M blended).
  • Muse Spark 1.1 streams 3.6× faster (194 vs 54 tokens per second).
MetricMuse Spark 1.1Qwen3.8-Flash-Next
BenchLeader Index65.059.1
Agents & tools score62.749.4
Coding score68.170.8
Composite score72.367.3
Human preference score66.3
Instruction following score77.8
Knowledge score73.259.5
Long context score65.466.4
Maths score52.2
Multimodal score65.064.3
Reasoning score69.2
Blended price $/M$2.00$0.230
Output speed194 tok/s54 tok/s
Time to first answer2.6 s39.9 s
Context window1.0M256k
SimpleQA Verified57.8%
SciCode58.2%
APEX-Agents41.9%
ProofBench39.0%
Epoch Capabilities Index154.6
LMArena Text1492
LMArena Hard Prompts1511
LMArena Coding1531
LMArena WebDev15411631
LMArena Vision1293
LMArena Agent-2.60.8
LiveBench76.2%
LiveBench Reasoning87.4%
LiveBench Coding72.5%
LiveBench Agentic Coding61.6%
LiveBench Mathematics85.8%
LiveBench Data Analysis74.2%
LiveBench Language74.6%
AA Intelligence Index34.339.9
AA-LCR77.7%79.7%
MMMU-Pro79.8%
AA-Omniscience28.1-9.7
GPQA Diamond (AA)89.8%92.3%
Humanity's Last Exam (AA)46.2%38.0%
SciCode (AA)58.8%50.6%
SWE-Bench Pro61.5%
MCP Atlas88.1%
MultiChallenge75.3%
PRBench Finance55.0%
PRBench Legal57.0%
MultiNRC65.6%
EQ-Bench 41260

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Muse Spark 1.1 vs Qwen3.8-Flash-Next: questions

Is Muse Spark 1.1 better than Qwen3.8-Flash-Next?
Muse Spark 1.1 leads on quality: 65.0 vs 59.1. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.1 is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Muse Spark 1.1 better than Qwen3.8-Flash-Next for coding?
Qwen3.8-Flash-Next scores higher in coding (71 vs 68 on the category index, where 50 is average).
Is Muse Spark 1.1 better than Qwen3.8-Flash-Next for agentic tasks?
Muse Spark 1.1 scores higher in agentic tasks (63 vs 49 on the category index, where 50 is average).
Which is cheaper, Muse Spark 1.1 or Qwen3.8-Flash-Next?
Qwen3.8-Flash-Next is cheaper: $0.230 against $2.00 per million tokens, blended at three input tokens per output token.
Which is faster, Muse Spark 1.1 or Qwen3.8-Flash-Next?
Muse Spark 1.1 streams faster: 194 against 54 output tokens per second.
Which has the larger context window?
Muse Spark 1.1 accepts more context: 1.0M against 256k tokens.