BenchLeader

Muse Spark 1.3 vs Qwen3 7

Verdict
  • Muse Spark 1.3 (xhigh) leads on quality: 63.9 vs 61.3.
  • Muse Spark 1.3 (xhigh) is stronger in agents & tools, coding, composite, knowledge, long context, multimodal.
  • Qwen3 7 (max) is stronger in human preference, instruction following, maths, reasoning.
  • Muse Spark 1.3 (xhigh) is 1.9× cheaper ($2.00 vs $3.75 per 1M blended).
MetricMuse Spark 1.3 (xhigh)Qwen3 7 (max)
BenchLeader Index63.961.3
Agents & tools score60.256.0
Coding score70.361.4
Composite score78.656.4
Knowledge score75.163.6
Long context score68.166.0
Multimodal score66.6
Human preference score57.2
Instruction following score77.6
Maths score57.4
Reasoning score66.8
Blended price $/M$2.00$3.75
Output speed190 tok/s169 tok/s
Time to first answer35.7 s16.5 s
Context window1M1M
GPQA Diamond90.9%
FrontierMath Tiers 1–364.6%
FrontierMath Tier 434.1%
OTIS Mock AIME95.6%
SWE-bench Verified (Epoch)77.3%
SimpleQA Verified55.8%
SimpleBench70.4%
SciCode48.8%
ProofBench26.0%
Epoch Capabilities Index153.7
LMArena Text1474
LMArena Hard Prompts1495
LMArena Coding1525
LMArena WebDev16251517
LMArena Agent-3.1
LiveBench81.6%73.1%
LiveBench Reasoning89.7%83.3%
LiveBench Coding81.1%74.2%
LiveBench Agentic Coding64.1%43.6%
LiveBench Mathematics96.0%85.3%
LiveBench Data Analysis79.6%71.8%
LiveBench Language82.8%79.7%
AA Intelligence Index45.229.9
IFBench80.5%
AA-LCR83.0%79.0%
MMMU-Pro82.0%
AA-Omniscience23.113.5
Terminal-Bench Hard50.8%
GPQA Diamond (AA)94.1%92.3%
Humanity's Last Exam (AA)47.5%40.5%
SciCode (AA)59.7%49.5%
τ²-Bench Telecom (AA)94.7%
LiveCodeBench87.1%
MMLU-Pro89.3%
IOI46.8%
LegalBench84.9%
CorpFin63.7%
TaxEval75.3%
Terminal-Bench 2.1 (Vals)72.3%61.0%
SWE-bench (Vals)68.8%
GPQA Diamond (Vals)90.2%
Vals Index60.344.8
EQ-Bench 41110

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Muse Spark 1.3 vs Qwen3 7: questions

Is Muse Spark 1.3 better than Qwen3 7?
Muse Spark 1.3 (xhigh) leads on quality: 63.9 vs 61.3. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 (xhigh) is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Muse Spark 1.3 better than Qwen3 7 for coding?
Muse Spark 1.3 scores higher in coding (70 vs 61 on the category index, where 50 is average).
Is Muse Spark 1.3 better than Qwen3 7 for agentic tasks?
Muse Spark 1.3 scores higher in agentic tasks (60 vs 56 on the category index, where 50 is average).
Which is cheaper, Muse Spark 1.3 or Qwen3 7?
Muse Spark 1.3 is cheaper: $2.00 against $3.75 per million tokens, blended at three input tokens per output token.
Which is faster, Muse Spark 1.3 or Qwen3 7?
Muse Spark 1.3 streams faster: 190 against 169 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.