BenchLeader

Muse Spark 1.3 vs Qwen3.5 397B A17B

Verdict
  • Muse Spark 1.3 leads on quality: 68.3 vs 58.8.
  • Muse Spark 1.3 is stronger in composite, knowledge, long context.
  • Qwen3.5 397B A17B is stronger in agents & tools, coding, human preference, instruction following, maths, multimodal, reasoning.
  • Qwen3.5 397B A17B is 5.2× cheaper ($0.387 vs $2.00 per 1M blended).
  • Muse Spark 1.3 streams 2.8× faster (241 vs 85 tokens per second).
MetricMuse Spark 1.3Qwen3.5 397B A17B
BenchLeader Index68.358.8
Composite score90.353.4
Knowledge score79.149.7
Long context score68.265.3
Agents & tools score69.7
Coding score51.6
Human preference score63.5
Instruction following score76.6
Maths score50.2
Multimodal score61.5
Reasoning score63.0
Blended price $/M$2.00$0.387
Output speed241 tok/s85 tok/s
Time to first answer30.5 s39.8 s
Context window1.0M262k
GPQA Diamond85.9%
FrontierMath Tiers 1–329.5%
OTIS Mock AIME88.9%
Epoch Capabilities Index147.0
LMArena Text1442
LMArena Hard Prompts1464
LMArena Coding1491
LMArena WebDev1399
LMArena Vision1265
AA Intelligence Index48.219.1
IFBench78.8%
AA-LCR83.0%77.3%
MMMU-Pro77.3%
AA-Omniscience25-30.8
Terminal-Bench Hard40.9%
GPQA Diamond (AA)93.5%89.3%
Humanity's Last Exam (AA)48.7%29.0%
SciCode (AA)58.8%44.8%
τ²-Bench Telecom (AA)95.6%
AIME 202694.2%
HMMT February 202687.9%
PRBench Finance59.5%
PRBench Legal61.6%

Data as of 2026-09-14. Best configuration of each model; every score links to its source on the model pages.

Muse Spark 1.3 vs Qwen3.5 397B A17B: questions

Is Muse Spark 1.3 better than Qwen3.5 397B A17B?
Muse Spark 1.3 leads on quality: 68.3 vs 58.8. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 is ahead overall as of 2026-09-14, but check the category scores for your use.
Which is cheaper, Muse Spark 1.3 or Qwen3.5 397B A17B?
Qwen3.5 397B A17B is cheaper: $0.387 against $2.00 per million tokens, blended at three input tokens per output token.
Which is faster, Muse Spark 1.3 or Qwen3.5 397B A17B?
Muse Spark 1.3 streams faster: 241 against 85 output tokens per second.
Which has the larger context window?
Muse Spark 1.3 accepts more context: 1.0M against 262k tokens.