BenchLeader

Muse Spark 1.1 vs Qwen3.6 Plus

Verdict
  • Muse Spark 1.1 leads on quality: 65.0 vs 59.2.
  • Muse Spark 1.1 is stronger in agents & tools, coding, composite, human preference, instruction following, knowledge, multimodal, reasoning.
  • Qwen3.6 Plus is stronger in long context, maths.
  • Qwen3.6 Plus is 1.8× cheaper ($1.13 vs $2.00 per 1M blended).
  • Muse Spark 1.1 streams 3.4× faster (194 vs 56 tokens per second).
MetricMuse Spark 1.1Qwen3.6 Plus
BenchLeader Index65.059.2
Agents & tools score62.757.0
Coding score68.152.8
Composite score72.348.3
Human preference score66.363.9
Instruction following score77.873.0
Knowledge score73.259.0
Long context score65.465.7
Maths score52.256.7
Multimodal score65.062.4
Reasoning score69.264.3
Blended price $/M$2.00$1.13
Output speed194 tok/s56 tok/s
Time to first answer2.6 s101.1 s
Context window1.0M1M
GPQA Diamond88.4%
FrontierMath Tiers 1–338.3%
OTIS Mock AIME93.3%
SWE-bench Verified (Epoch)57.9%
SimpleQA Verified57.8%44.1%
SciCode58.2%40.7%
APEX-Agents41.9%
ProofBench39.0%
Epoch Capabilities Index154.6147.7
LMArena Text14921444
LMArena Hard Prompts15111468
LMArena Coding15311495
LMArena WebDev15411460
LMArena Vision1293
LMArena Agent-2.6
LiveBench68.9%
LiveBench Reasoning75.8%
LiveBench Coding78.2%
LiveBench Agentic Coding41.4%
LiveBench Mathematics83.7%
LiveBench Data Analysis69.9%
LiveBench Language75.0%
AA Intelligence Index34.327.0
IFBench75.2%
AA-LCR77.7%78.3%
MMMU-Pro78.0%
AA-Omniscience28.10.9
Terminal-Bench Hard43.9%
GPQA Diamond (AA)89.8%88.2%
Humanity's Last Exam (AA)46.2%27.9%
SciCode (AA)58.8%
τ²-Bench Telecom (AA)97.7%
AIME (Vals)94.6%
LiveCodeBench86.0%
MMLU-Pro87.7%
LegalBench84.2%
CorpFin61.9%
TaxEval74.7%
Terminal-Bench 2.1 (Vals)53.2%
SWE-bench (Vals)73.4%
GPQA Diamond (Vals)87.4%
Vals Index32.0
SWE-Bench Pro61.5%
MCP Atlas88.1%
MultiChallenge75.3%
PRBench Finance55.0%
PRBench Legal57.0%
MultiNRC65.6%
EQ-Bench 41260

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Muse Spark 1.1 vs Qwen3.6 Plus: questions

Is Muse Spark 1.1 better than Qwen3.6 Plus?
Muse Spark 1.1 leads on quality: 65.0 vs 59.2. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.1 is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Muse Spark 1.1 better than Qwen3.6 Plus for coding?
Muse Spark 1.1 scores higher in coding (68 vs 53 on the category index, where 50 is average).
Is Muse Spark 1.1 better than Qwen3.6 Plus for agentic tasks?
Muse Spark 1.1 scores higher in agentic tasks (63 vs 57 on the category index, where 50 is average).
Which is cheaper, Muse Spark 1.1 or Qwen3.6 Plus?
Qwen3.6 Plus is cheaper: $1.13 against $2.00 per million tokens, blended at three input tokens per output token.
Which is faster, Muse Spark 1.1 or Qwen3.6 Plus?
Muse Spark 1.1 streams faster: 194 against 56 output tokens per second.
Which has the larger context window?
Muse Spark 1.1 accepts more context: 1.0M against 1M tokens.