BenchLeader

Muse Spark 1.1 vs Qwen3.5 397B A17B

Verdict
  • Muse Spark 1.1 leads on quality: 65.0 vs 58.7.
  • Muse Spark 1.1 is stronger in coding, composite, human preference, instruction following, knowledge, long context, maths, multimodal, reasoning.
  • Qwen3.5 397B A17B is stronger in agents & tools.
  • Qwen3.5 397B A17B is 5.2× cheaper ($0.387 vs $2.00 per 1M blended).
  • Muse Spark 1.1 streams 2.5× faster (194 vs 78 tokens per second).
MetricMuse Spark 1.1Qwen3.5 397B A17B
BenchLeader Index65.058.7
Agents & tools score62.769.4
Coding score68.151.7
Composite score72.353.1
Human preference score66.363.6
Instruction following score77.876.1
Knowledge score73.249.5
Long context score65.465.2
Maths score52.250.1
Multimodal score65.061.6
Reasoning score69.263.0
Blended price $/M$2.00$0.387
Output speed194 tok/s78 tok/s
Time to first answer2.6 s42.7 s
Context window1.0M262k
GPQA Diamond85.9%
FrontierMath Tiers 1–329.5%
OTIS Mock AIME88.9%
SimpleQA Verified57.8%
SciCode58.2%
APEX-Agents41.9%
ProofBench39.0%
Epoch Capabilities Index154.6147.0
LMArena Text14921441
LMArena Hard Prompts15111463
LMArena Coding15311491
LMArena WebDev15411399
LMArena Vision12931265
LMArena Agent-2.6
AA Intelligence Index34.319.1
IFBench78.8%
AA-LCR77.7%77.3%
MMMU-Pro77.3%
AA-Omniscience28.1-30.8
Terminal-Bench Hard40.9%
GPQA Diamond (AA)89.8%89.3%
Humanity's Last Exam (AA)46.2%29.0%
SciCode (AA)58.8%44.8%
τ²-Bench Telecom (AA)95.6%
AIME 202694.2%
HMMT February 202687.9%
SWE-Bench Pro61.5%
MCP Atlas88.1%
MultiChallenge75.3%
PRBench Finance55.0%
PRBench Legal57.0%
MultiNRC65.6%
EQ-Bench 41260

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Muse Spark 1.1 vs Qwen3.5 397B A17B: questions

Is Muse Spark 1.1 better than Qwen3.5 397B A17B?
Muse Spark 1.1 leads on quality: 65.0 vs 58.7. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.1 is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Muse Spark 1.1 better than Qwen3.5 397B A17B for coding?
Muse Spark 1.1 scores higher in coding (68 vs 52 on the category index, where 50 is average).
Is Muse Spark 1.1 better than Qwen3.5 397B A17B for agentic tasks?
Qwen3.5 397B A17B scores higher in agentic tasks (69 vs 63 on the category index, where 50 is average).
Which is cheaper, Muse Spark 1.1 or Qwen3.5 397B A17B?
Qwen3.5 397B A17B is cheaper: $0.387 against $2.00 per million tokens, blended at three input tokens per output token.
Which is faster, Muse Spark 1.1 or Qwen3.5 397B A17B?
Muse Spark 1.1 streams faster: 194 against 78 output tokens per second.
Which has the larger context window?
Muse Spark 1.1 accepts more context: 1.0M against 262k tokens.