BenchLeader

Muse Spark 1.1 vs Qwen3.8 Max

Verdict
  • Qwen3.8 Max (max) leads on quality: 64.2 vs 63.1.
  • Muse Spark 1.1 is stronger in coding, instruction following, knowledge.
  • Qwen3.8 Max (max) is stronger in agents & tools, human preference, maths, multimodal, reasoning, composite, long context.
  • Muse Spark 1.1 is 1.5× cheaper ($2.00 vs $3.00 per 1M blended).
  • Muse Spark 1.1 streams 5.5× faster (201 vs 37 tokens per second).
MetricMuse Spark 1.1Qwen3.8 Max (max)
BenchLeader Index63.164.2
Agents & tools score51.558.2
Coding score66.664.6
Human preference score65.768.0
Instruction following score74.2–
Knowledge score70.362.7
Maths score61.965.1
Multimodal score63.866.2
Reasoning score68.472.2
Composite score–70.7
Long context score–65.4
Blended price $/M$2.00$3.00
Output speed201 tok/s37 tok/s
Time to first answer1.9 s58.9 s
Context window1.0M1M
SimpleQA Verified57.8%–
Terminal-Bench–27.0%
SciCode58.8%53.2%
APEX-Agents31.8%63.3%
ProofBench39.0%58.0%
Epoch Capabilities Index154.2156.4
LMArena Text14911483
LMArena Hard Prompts15111504
LMArena Coding15331524
LMArena WebDev15421672
LMArena Vision12931314
LMArena Agent-52.3
LiveBench–78.5%
LiveBench Reasoning–88.2%
LiveBench Coding–72.9%
LiveBench Agentic Coding–64.7%
LiveBench Mathematics–91.3%
LiveBench Data Analysis–78.4%
LiveBench Language–79.7%
LiveBench Instruction Following–74.1%
AA Intelligence Index v4.3.2–45.4
AA-LCR–80.3%
MMMU-Pro–82.8%
AA-Omniscience–12.0
GPQA Diamond (AA)–92.8%
Humanity's Last Exam (AA)–43.1%
SciCode (AA)–53.2%
LiveCodeBench–87.8%
MMLU-Pro–88.6%
IOI–68.9%
LegalBench–83.6%
CorpFin–65.8%
TaxEval–75.5%
Terminal-Bench 2.1 (Vals)–67.4%
SWE-bench (Vals)–85.6%
GPQA Diamond (Vals)–93.7%
Vals Index–48.3
SWE-Bench Pro61.5%–
MCP Atlas88.1%–
MultiChallenge75.3%–
PRBench Finance55.0%–
PRBench Legal57.0%–
MultiNRC65.6%–
EQ-Bench 41260–
CritPt–20.0%
GDPval-AA v2.1–58.6%
τ³-Banking (AA)–51.3%
ITBench SRE (AA)–40.3%
Analyst Agent (AA)–45.0%
APEX-Agents (AA)–42.4%
Code Migration–24.0%
CyberBench–28.6%
Excel Modeling Benchmark–60.1%
Finance Agent v2–50.6%
Harvey's Legal Agent Benchmark–10.4%
Legal Research Bench–47.6%
MedCode–40.7%
MedScribe–85.0%
MMMU-Pro (Vals)–88.0%
MortgageTax–64.0%
MysteryMechanism–23.9%
ProgramBench–0.0%
Public Benefits Bench–67.1%
SAGE–51.3%
SkillsBench–42.0%
Tax Agent Bench–66.0%
Terminal-Bench 4.0 (Vals)–34.3%
Terminal-Bench Science–1.4%
Vals Multimodal Index–65.4%
Vibe Code Bench 1-100–12.8%
Vibe Code Bench v1.1–64.7%
FORTRESS12.4%–
SWE-Bench Pro (private)61.5%–
LMArena Maths14951497
LMArena Creative Writing14451470
LMArena Instruction Following14721474
LMArena Multi-turn14971492
LMArena Longer Queries14811492
LMArena Document1480–
DeepSWE v1.153.3%–
GBAEval7.9%–
Vending-Bench 26520.5–
GDP.pdf15.0%–
Terminal-Bench 4.0 (AA)–38.9%
Terminal-Bench 2.1 (AA)–88.8%
AutomationBench–56.2%
GDP.pdf–22.8%
MLCR–20.0%
EnterpriseOps-Gym–47.6%
AA-Omniscience: accuracy–31.9%
AA-Omniscience: non-hallucination–71.2%
AA-Briefcase v1.1–1617

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Muse Spark 1.1 vs Qwen3.8 Max: questions

Is Muse Spark 1.1 better than Qwen3.8 Max?
Qwen3.8 Max (max) leads on quality: 64.2 vs 63.1. The BenchLeader Index combines every independent quality benchmark; Qwen3.8 Max (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Muse Spark 1.1 better than Qwen3.8 Max for coding?
Muse Spark 1.1 scores higher in coding (67 vs 65 on the category index, where 50 is average).
Is Muse Spark 1.1 better than Qwen3.8 Max for agentic tasks?
Qwen3.8 Max scores higher in agentic tasks (58 vs 52 on the category index, where 50 is average).
Which is cheaper, Muse Spark 1.1 or Qwen3.8 Max?
Muse Spark 1.1 is cheaper: $2.00 against $3.00 per million tokens, blended at three input tokens per output token.
Which is faster, Muse Spark 1.1 or Qwen3.8 Max?
Muse Spark 1.1 streams faster: 201 against 37 output tokens per second.
Which has the larger context window?
Muse Spark 1.1 accepts more context: 1.0M against 1M tokens.