BenchLeader

Muse Spark 1.2 vs Qwen3.8 Max

Verdict
  • Qwen3.8 Max (max) leads on quality: 64.2 vs 62.7.
  • Muse Spark 1.2 (xhigh) is stronger in human preference, knowledge, maths.
  • Qwen3.8 Max (max) is stronger in agents & tools, coding, composite, long context, multimodal, reasoning.
  • Muse Spark 1.2 (xhigh) is 1.5× cheaper ($2.00 vs $3.00 per 1M blended).
  • Muse Spark 1.2 (xhigh) streams 8.8× faster (325 vs 37 tokens per second).
MetricMuse Spark 1.2 (xhigh)Qwen3.8 Max (max)
BenchLeader Index62.764.2
Agents & tools score55.958.2
Coding score55.064.6
Composite score66.670.7
Human preference score69.168.0
Knowledge score66.962.7
Long context score64.765.4
Maths score66.765.1
Multimodal score65.366.2
Reasoning score69.872.2
Blended price $/M$2.00$3.00
Output speed325 tok/s37 tok/s
Time to first answer18.2 s58.9 s
Context window1.0M1M
SimpleQA Verified60.3%–
Terminal-Bench–27.0%
SciCode56.4%53.2%
WeirdML60.3%–
APEX-Agents–63.3%
ProofBench–58.0%
Epoch Capabilities Index–156.4
LMArena Text14921483
LMArena Hard Prompts15081504
LMArena Coding15311524
LMArena WebDev15331672
LMArena Vision13051314
LMArena Agent-3.72.3
LiveBench78.0%78.5%
LiveBench Reasoning90.0%88.2%
LiveBench Coding77.5%72.9%
LiveBench Agentic Coding57.6%64.7%
LiveBench Mathematics91.2%91.3%
LiveBench Data Analysis76.5%78.4%
LiveBench Language78.6%79.7%
LiveBench Instruction Following74.3%74.1%
AA Intelligence Index v4.3.239.645.4
AA-LCR79.0%80.3%
MMMU-Pro–82.8%
AA-Omniscience27.212.0
GPQA Diamond (AA)90.4%92.8%
Humanity's Last Exam (AA)45.5%43.1%
SciCode (AA)57.4%53.2%
LiveCodeBench–87.8%
MMLU-Pro88.3%88.6%
IOI21.8%68.9%
LegalBench85.3%83.6%
CorpFin70.9%65.8%
TaxEval80.4%75.5%
Terminal-Bench 2.1 (Vals)69.7%67.4%
SWE-bench (Vals)86.6%85.6%
GPQA Diamond (Vals)–93.7%
Vals Index49.348.3
CritPt17.7%20.0%
GDPval-AA v2.148.9%58.6%
τ³-Banking (AA)34.9%51.3%
ITBench SRE (AA)–40.3%
Analyst Agent (AA)–45.0%
APEX-Agents (AA)–42.4%
BioMysteryBench64.8%–
Code Migration29.9%24.0%
CyberBench69.5%28.6%
Excel Modeling Benchmark57.0%60.1%
Finance Agent v260.6%50.6%
Harvey's Legal Agent Benchmark25.4%10.4%
Legal Research Bench43.8%47.6%
MedCode49.4%40.7%
MedScribe90.1%85.0%
MMMU-Pro (Vals)86.1%88.0%
MortgageTax65.4%64.0%
MysteryMechanism–23.9%
ProgramBench–0.0%
Public Benefits Bench68.5%67.1%
SAGE47.7%51.3%
SkillsBench53.0%42.0%
Tax Agent Bench56.9%66.0%
Terminal-Bench 4.0 (Vals)6.1%34.3%
Terminal-Bench Science–1.4%
Vals Multimodal Index–65.4%
Vibe Code Bench 1-100–12.8%
Vibe Code Bench v1.179.1%64.7%
LMArena Maths14811497
LMArena Creative Writing14541470
LMArena Instruction Following14711474
LMArena Multi-turn15061492
LMArena Longer Queries14891492
DeepSWE v1.154.9%–
LMCA48.4%–
DTBench94.7%–
GDP.pdf16.0%–
FrontierSWE12.0%–
Terminal-Bench 4.0 (AA)7.1%38.9%
Terminal-Bench 2.1 (AA)80.2%88.8%
AutomationBench–56.2%
GDP.pdf–22.8%
MLCR–20.0%
EnterpriseOps-Gym–47.6%
AA-Omniscience: accuracy45.4%31.9%
AA-Omniscience: non-hallucination66.7%71.2%
AA-Briefcase v1.1–1617

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Muse Spark 1.2 vs Qwen3.8 Max: questions

Is Muse Spark 1.2 better than Qwen3.8 Max?
Qwen3.8 Max (max) leads on quality: 64.2 vs 62.7. The BenchLeader Index combines every independent quality benchmark; Qwen3.8 Max (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Muse Spark 1.2 better than Qwen3.8 Max for coding?
Qwen3.8 Max scores higher in coding (65 vs 55 on the category index, where 50 is average).
Is Muse Spark 1.2 better than Qwen3.8 Max for agentic tasks?
Qwen3.8 Max scores higher in agentic tasks (58 vs 56 on the category index, where 50 is average).
Which is cheaper, Muse Spark 1.2 or Qwen3.8 Max?
Muse Spark 1.2 is cheaper: $2.00 against $3.00 per million tokens, blended at three input tokens per output token.
Which is faster, Muse Spark 1.2 or Qwen3.8 Max?
Muse Spark 1.2 streams faster: 325 against 37 output tokens per second.
Which has the larger context window?
Muse Spark 1.2 accepts more context: 1.0M against 1M tokens.