BenchLeader

Muse Spark 1.3 vs Qwen3.7 Plus

Verdict
  • Muse Spark 1.3 leads on quality: 68.3 vs 59.8.
  • Muse Spark 1.3 is stronger in composite, knowledge, long context.
  • Qwen3.7 Plus is stronger in agents & tools, coding, human preference, instruction following, maths, multimodal, reasoning.
  • Qwen3.7 Plus is 1.8× cheaper ($1.13 vs $2.00 per 1M blended).
  • Muse Spark 1.3 streams 3.6× faster (241 vs 66 tokens per second).
MetricMuse Spark 1.3Qwen3.7 Plus
BenchLeader Index68.359.8
Composite score90.361.9
Knowledge score79.164.9
Long context score68.263.1
Agents & tools score44.7
Coding score46.7
Human preference score65.2
Instruction following score75.9
Maths score65.2
Multimodal score64.0
Reasoning score64.2
Blended price $/M$2.00$1.13
Output speed241 tok/s66 tok/s
Time to first answer30.5 s32.4 s
Context window1.0M1M
GPQA Diamond87.9%
OTIS Mock AIME93.3%
OSWorld-Verified 2.02.8%
SciCode45.5%
FrontierCode10.2%
Epoch Capabilities Index147.4
LMArena Text1456
LMArena Hard Prompts1473
LMArena Coding1501
LMArena Vision1278
LMArena Agent-5.4
AA Intelligence Index48.225.8
IFBench78.0%
AA-LCR83.0%73.0%
MMMU-Pro80.5%
AA-Omniscience251.1
Terminal-Bench Hard47.0%
GPQA Diamond (AA)93.5%90.0%
Humanity's Last Exam (AA)48.7%35.6%
SciCode (AA)58.8%46.1%
τ²-Bench Telecom (AA)93.0%
Terminal-Bench 2.1 (Vals)52.8%
Vals Index38.6
PRBench Finance59.5%
PRBench Legal61.6%

Data as of 2026-09-14. Best configuration of each model; every score links to its source on the model pages.

Muse Spark 1.3 vs Qwen3.7 Plus: questions

Is Muse Spark 1.3 better than Qwen3.7 Plus?
Muse Spark 1.3 leads on quality: 68.3 vs 59.8. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 is ahead overall as of 2026-09-14, but check the category scores for your use.
Which is cheaper, Muse Spark 1.3 or Qwen3.7 Plus?
Qwen3.7 Plus is cheaper: $1.13 against $2.00 per million tokens, blended at three input tokens per output token.
Which is faster, Muse Spark 1.3 or Qwen3.7 Plus?
Muse Spark 1.3 streams faster: 241 against 66 output tokens per second.
Which has the larger context window?
Muse Spark 1.3 accepts more context: 1.0M against 1M tokens.