BenchLeader

Muse Spark vs Qwen3.5 397B A17B

Verdict
  • Muse Spark leads on quality: 65.6 vs 58.7.
  • Muse Spark is stronger in coding, composite, human preference, instruction following, knowledge, long context, maths, multimodal, reasoning.
  • Qwen3.5 397B A17B is stronger in agents & tools.
MetricMuse SparkQwen3.5 397B A17B
BenchLeader Index65.658.7
Agents & tools score66.469.4
Coding score66.251.7
Composite score68.653.1
Human preference score69.363.6
Instruction following score79.776.1
Knowledge score64.649.5
Long context score65.565.2
Maths score56.450.1
Multimodal score65.961.6
Reasoning score70.263.0
Blended price $/M$0.387
Output speed78 tok/s
Time to first answer42.7 s
Context window262k262k
GPQA Diamond89.8%85.9%
FrontierMath Tiers 1–329.5%
OTIS Mock AIME88.9%88.9%
Humanity's Last Exam40.6%
SciCode51.5%
ProofBench17.0%
Epoch Capabilities Index152.1147.0
LMArena Text14881441
LMArena Hard Prompts15051463
LMArena Coding15261491
LMArena WebDev1399
LMArena Vision13061265
AA Intelligence Index31.319.1
IFBench75.9%78.8%
AA-LCR78.0%77.3%
MMMU-Pro80.5%77.3%
AA-Omniscience7.2-30.8
Terminal-Bench Hard45.5%40.9%
GPQA Diamond (AA)88.4%89.3%
Humanity's Last Exam (AA)40.7%29.0%
SciCode (AA)44.8%
τ²-Bench Telecom (AA)91.5%95.6%
AIME (Vals)96.9%
MMLU-Pro87.3%
LegalBench84.2%
CorpFin65.1%
TaxEval77.7%
SWE-bench (Vals)74.4%
GPQA Diamond (Vals)89.7%
AIME 202694.2%
HMMT February 202687.9%
SWE-Bench Pro55.0%
MCP Atlas82.2%
MultiChallenge75.5%
PRBench Finance52.4%
PRBench Legal52.3%
MultiNRC59.0%
TutorBench68.5%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Muse Spark vs Qwen3.5 397B A17B: questions

Is Muse Spark better than Qwen3.5 397B A17B?
Muse Spark leads on quality: 65.6 vs 58.7. The BenchLeader Index combines every independent quality benchmark; Muse Spark is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Muse Spark better than Qwen3.5 397B A17B for coding?
Muse Spark scores higher in coding (66 vs 52 on the category index, where 50 is average).
Is Muse Spark better than Qwen3.5 397B A17B for agentic tasks?
Qwen3.5 397B A17B scores higher in agentic tasks (69 vs 66 on the category index, where 50 is average).
Which has the larger context window?
Both accept 262k tokens of context.