BenchLeader

Muse Spark vs Qwen3.8 Max

Verdict
  • Muse Spark and Qwen3.8 Max (max) are level on quality (64.6 vs 64.2).
  • Muse Spark is stronger in agents & tools, coding, human preference, instruction following, knowledge.
  • Qwen3.8 Max (max) is stronger in composite, long context, maths, multimodal, reasoning.
MetricMuse SparkQwen3.8 Max (max)
BenchLeader Index64.664.2
Agents & tools score66.058.2
Coding score65.364.6
Composite score65.070.7
Human preference score68.768.0
Instruction following score78.3–
Knowledge score63.762.7
Long context score64.265.4
Maths score59.365.1
Multimodal score64.666.2
Reasoning score65.972.2
Blended price $/M–$3.00
Output speed–37 tok/s
Time to first answer–58.9 s
Context window262k1M
GPQA Diamond89.8%–
OTIS Mock AIME88.9%–
Humanity's Last Exam40.6%–
Terminal-Bench–27.0%
SciCode51.5%53.2%
APEX-Agents–63.3%
ProofBench17.0%58.0%
Epoch Capabilities Index152.0156.4
LMArena Text14891483
LMArena Hard Prompts15061504
LMArena Coding15291524
LMArena WebDev–1672
LMArena Vision13061314
LMArena Agent–2.3
LiveBench–78.5%
LiveBench Reasoning–88.2%
LiveBench Coding–72.9%
LiveBench Agentic Coding–64.7%
LiveBench Mathematics–91.3%
LiveBench Data Analysis–78.4%
LiveBench Language–79.7%
LiveBench Instruction Following–74.1%
AA Intelligence Index v4.3.231.345.4
IFBench75.9%–
AA-LCR78.0%80.3%
MMMU-Pro80.5%82.8%
AA-Omniscience7.212.0
Terminal-Bench Hard45.5%–
GPQA Diamond (AA)88.4%92.8%
Humanity's Last Exam (AA)40.7%43.1%
SciCode (AA)–53.2%
τ²-Bench Telecom (AA)91.5%–
AIME (Vals)96.9%–
LiveCodeBench–87.8%
MMLU-Pro87.3%88.6%
IOI–68.9%
LegalBench84.2%83.6%
CorpFin65.1%65.8%
TaxEval77.7%75.5%
Terminal-Bench 2.1 (Vals)–67.4%
SWE-bench (Vals)74.4%85.6%
GPQA Diamond (Vals)89.7%93.7%
Vals Index–48.3
SWE-Bench Pro55.0%–
MCP Atlas82.2%–
MultiChallenge75.5%–
PRBench Finance52.4%–
PRBench Legal52.3%–
MultiNRC59.0%–
TutorBench68.5%–
CritPt11.3%20.0%
GDPval-AA v2.125.1%58.6%
τ³-Banking (AA)–51.3%
ITBench SRE (AA)–40.3%
Analyst Agent (AA)–45.0%
APEX-Agents (AA)–42.4%
CaseLaw v263.1%–
Code Migration–24.0%
CyberBench–28.6%
Excel Modeling Benchmark–60.1%
Finance Agent v2–50.6%
Harvey's Legal Agent Benchmark–10.4%
Legal Research Bench–47.6%
MedCode51.3%40.7%
MedScribe85.9%85.0%
MMMU-Pro (Vals)87.4%88.0%
MortgageTax–64.0%
MysteryMechanism–23.9%
ProgramBench–0.0%
Public Benefits Bench–67.1%
SAGE–51.3%
SkillsBench–42.0%
Tax Agent Bench–66.0%
Terminal-Bench 2.0 (Vals)59.5%–
Terminal-Bench 4.0 (Vals)–34.3%
Terminal-Bench Science–1.4%
Vals Multimodal Index–65.4%
Vibe Code Bench 1-100–12.8%
Vibe Code Bench v1.119.7%64.7%
FORTRESS20.2%–
SWE-Bench Pro (private)55.0%–
SWE Atlas: Codebase QnA24.2%–
SWE Atlas: Test Writing31.1%–
LMArena Maths14661497
LMArena Creative Writing14651470
LMArena Instruction Following14641474
LMArena Multi-turn14931492
LMArena Longer Queries14761492
LMArena Document1467–
Terminal-Bench 4.0 (AA)–38.9%
Terminal-Bench 2.1 (AA)62.2%88.8%
AutomationBench–56.2%
GDP.pdf–22.8%
MLCR–20.0%
EnterpriseOps-Gym–47.6%
AA-Omniscience: accuracy49.6%31.9%
AA-Omniscience: non-hallucination15.8%71.2%
AA-Briefcase v1.1–1617

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Muse Spark vs Qwen3.8 Max: questions

Is Muse Spark better than Qwen3.8 Max?
Muse Spark and Qwen3.8 Max (max) are level on quality (64.6 vs 64.2). The BenchLeader Index combines every independent quality benchmark; Muse Spark is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Muse Spark better than Qwen3.8 Max for coding?
Muse Spark scores higher in coding (65 vs 65 on the category index, where 50 is average).
Is Muse Spark better than Qwen3.8 Max for agentic tasks?
Muse Spark scores higher in agentic tasks (66 vs 58 on the category index, where 50 is average).
Which has the larger context window?
Qwen3.8 Max accepts more context: 1M against 262k tokens.