BenchLeader

Muse Spark 1.3 vs Qwen3.8 Max

Verdict
  • Muse Spark 1.3 (max) leads on quality: 67.7 vs 64.2.
  • Muse Spark 1.3 (max) is stronger in agents & tools, composite, human preference, knowledge, long context, reasoning.
  • Qwen3.8 Max (max) is stronger in maths, multimodal.
  • Muse Spark 1.3 (max) is 1.5× cheaper ($2.00 vs $3.00 per 1M blended).
  • Muse Spark 1.3 (max) streams 4.8× faster (175 vs 37 tokens per second).
MetricMuse Spark 1.3 (max)Qwen3.8 Max (max)
BenchLeader Index67.764.2
Agents & tools score70.358.2
Coding score64.664.6
Composite score84.170.7
Human preference score69.468.0
Knowledge score72.962.7
Long context score66.865.4
Maths score63.665.1
Multimodal score65.966.2
Reasoning score78.172.2
Blended price $/M$2.00$3.00
Output speed175 tok/s37 tok/s
Time to first answer70.5 s58.9 s
Context window1.0M1M
FrontierMath Tiers 1–374.0%–
FrontierMath Tier 446.3%–
Terminal-Bench–27.0%
SciCode58.8%53.2%
APEX-Agents–63.3%
ProofBench58.0%58.0%
Epoch Capabilities Index–156.4
LMArena Text14941483
LMArena Hard Prompts15171504
LMArena Coding15381524
LMArena WebDev16571672
LMArena Vision13091314
LMArena Agent42.3
LiveBench–78.5%
LiveBench Reasoning–88.2%
LiveBench Coding–72.9%
LiveBench Agentic Coding–64.7%
LiveBench Mathematics–91.3%
LiveBench Data Analysis–78.4%
LiveBench Language–79.7%
LiveBench Instruction Following–74.1%
AA Intelligence Index v4.3.248.145.4
AA-LCR83.0%80.3%
MMMU-Pro–82.8%
AA-Omniscience2512.0
GPQA Diamond (AA)93.5%92.8%
Humanity's Last Exam (AA)48.7%43.1%
SciCode (AA)58.8%53.2%
LiveCodeBench–87.8%
MMLU-Pro–88.6%
IOI56.6%68.9%
LegalBench–83.6%
CorpFin–65.8%
TaxEval–75.5%
Terminal-Bench 2.1 (Vals)79.0%67.4%
SWE-bench (Vals)–85.6%
GPQA Diamond (Vals)–93.7%
Vals Index58.248.3
CritPt24.9%20.0%
GDPval-AA v2.159.0%58.6%
τ³-Banking (AA)50.5%51.3%
ITBench SRE (AA)33.2%40.3%
Analyst Agent (AA)–45.0%
APEX-Agents (AA)–42.4%
Code Migration47.4%24.0%
CUA-bench5.8%–
CyberBench72.7%28.6%
Excel Modeling Benchmark67.4%60.1%
Finance Agent v260.0%50.6%
Harvey's Legal Agent Benchmark23.8%10.4%
Legal Research Bench55.3%47.6%
MedCode–40.7%
MedScribe–85.0%
MMMU-Pro (Vals)–88.0%
MortgageTax–64.0%
MysteryMechanism36.0%23.9%
ProgramBench–0.0%
Public Benefits Bench–67.1%
SAGE–51.3%
SkillsBench–42.0%
Tax Agent Bench72.4%66.0%
Terminal-Bench 4.0 (Vals)24.8%34.3%
Terminal-Bench Science10.0%1.4%
Vals Multimodal Index–65.4%
Vibe Code Bench 1-10020.5%12.8%
Vibe Code Bench v1.185.9%64.7%
LMArena Maths15021497
LMArena Creative Writing14561470
LMArena Instruction Following14831474
LMArena Multi-turn14871492
LMArena Longer Queries14971492
LMArena Document1468–
Chess Puzzles38.0%–
Mystery Game Puzzles25.0%–
CursorBench41.6%–
Terminal-Bench 4.0 (AA)33.3%38.9%
Terminal-Bench 2.1 (AA)84.3%88.8%
AutomationBench57.9%56.2%
GDP.pdf26.6%22.8%
MLCR43.3%20.0%
Harvey LAB8.9%–
EnterpriseOps-Gym–47.6%
AA-Omniscience: accuracy43.6%31.9%
AA-Omniscience: non-hallucination67.1%71.2%
AA-Briefcase v1.115811617

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Muse Spark 1.3 vs Qwen3.8 Max: questions

Is Muse Spark 1.3 better than Qwen3.8 Max?
Muse Spark 1.3 (max) leads on quality: 67.7 vs 64.2. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Muse Spark 1.3 better than Qwen3.8 Max for coding?
They are level in coding (65 each).
Is Muse Spark 1.3 better than Qwen3.8 Max for agentic tasks?
Muse Spark 1.3 scores higher in agentic tasks (70 vs 58 on the category index, where 50 is average).
Which is cheaper, Muse Spark 1.3 or Qwen3.8 Max?
Muse Spark 1.3 is cheaper: $2.00 against $3.00 per million tokens, blended at three input tokens per output token.
Which is faster, Muse Spark 1.3 or Qwen3.8 Max?
Muse Spark 1.3 streams faster: 175 against 37 output tokens per second.
Which has the larger context window?
Muse Spark 1.3 accepts more context: 1.0M against 1M tokens.