BenchLeader

Muse Spark 1.1 vs Qwen3 7

Verdict
  • Muse Spark 1.1 leads on quality: 65.0 vs 61.3.
  • Muse Spark 1.1 is stronger in agents & tools, coding, composite, human preference, instruction following, knowledge, multimodal, reasoning.
  • Qwen3 7 (max) is stronger in long context, maths.
  • Muse Spark 1.1 is 1.9× cheaper ($2.00 vs $3.75 per 1M blended).
MetricMuse Spark 1.1Qwen3 7 (max)
BenchLeader Index65.061.3
Agents & tools score62.756.0
Coding score68.161.4
Composite score72.356.4
Human preference score66.357.2
Instruction following score77.877.6
Knowledge score73.263.6
Long context score65.466.0
Maths score52.257.4
Multimodal score65.0
Reasoning score69.266.8
Blended price $/M$2.00$3.75
Output speed194 tok/s169 tok/s
Time to first answer2.6 s16.5 s
Context window1.0M1M
GPQA Diamond90.9%
FrontierMath Tiers 1–364.6%
FrontierMath Tier 434.1%
OTIS Mock AIME95.6%
SWE-bench Verified (Epoch)77.3%
SimpleQA Verified57.8%55.8%
SimpleBench70.4%
SciCode58.2%48.8%
APEX-Agents41.9%
ProofBench39.0%26.0%
Epoch Capabilities Index154.6153.7
LMArena Text14921474
LMArena Hard Prompts15111495
LMArena Coding15311525
LMArena WebDev15411517
LMArena Vision1293
LMArena Agent-2.6-3.1
LiveBench73.1%
LiveBench Reasoning83.3%
LiveBench Coding74.2%
LiveBench Agentic Coding43.6%
LiveBench Mathematics85.3%
LiveBench Data Analysis71.8%
LiveBench Language79.7%
AA Intelligence Index34.329.9
IFBench80.5%
AA-LCR77.7%79.0%
AA-Omniscience28.113.5
Terminal-Bench Hard50.8%
GPQA Diamond (AA)89.8%92.3%
Humanity's Last Exam (AA)46.2%40.5%
SciCode (AA)58.8%49.5%
τ²-Bench Telecom (AA)94.7%
LiveCodeBench87.1%
MMLU-Pro89.3%
IOI46.8%
LegalBench84.9%
CorpFin63.7%
TaxEval75.3%
Terminal-Bench 2.1 (Vals)61.0%
SWE-bench (Vals)68.8%
GPQA Diamond (Vals)90.2%
Vals Index44.8
SWE-Bench Pro61.5%
MCP Atlas88.1%
MultiChallenge75.3%
PRBench Finance55.0%
PRBench Legal57.0%
MultiNRC65.6%
EQ-Bench 412601110

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Muse Spark 1.1 vs Qwen3 7: questions

Is Muse Spark 1.1 better than Qwen3 7?
Muse Spark 1.1 leads on quality: 65.0 vs 61.3. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.1 is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Muse Spark 1.1 better than Qwen3 7 for coding?
Muse Spark 1.1 scores higher in coding (68 vs 61 on the category index, where 50 is average).
Is Muse Spark 1.1 better than Qwen3 7 for agentic tasks?
Muse Spark 1.1 scores higher in agentic tasks (63 vs 56 on the category index, where 50 is average).
Which is cheaper, Muse Spark 1.1 or Qwen3 7?
Muse Spark 1.1 is cheaper: $2.00 against $3.75 per million tokens, blended at three input tokens per output token.
Which is faster, Muse Spark 1.1 or Qwen3 7?
Muse Spark 1.1 streams faster: 194 against 169 output tokens per second.
Which has the larger context window?
Muse Spark 1.1 accepts more context: 1.0M against 1M tokens.