BenchLeader

Grok 4.20 Multi-Agent vs Muse Spark 1.3

Verdict
  • Muse Spark 1.3 (max) leads on quality: 69.4 vs 59.2.
  • Muse Spark 1.3 (max) is stronger in coding, human preference, maths, multimodal, reasoning, agents & tools, composite, knowledge, long context.
  • Grok 4.20 Multi-Agent is 1.3× cheaper ($1.56 vs $2.00 per 1M blended).
MetricGrok 4.20 Multi-AgentMuse Spark 1.3 (max)
BenchLeader Index59.269.4
Coding score65.566.7
Human preference score67.069.8
Maths score64.466.0
Multimodal score60.467.4
Reasoning score65.982.9
Agents & tools score71.0
Composite score89.7
Knowledge score75.9
Long context score68.0
Blended price $/M$1.56$2.00
Output speed304 tok/s241 tok/s
Time to first answer9.9 s29.3 s
Context window1M1.0M
FrontierMath Tiers 1–374.0%
FrontierMath Tier 446.3%
LMArena Text14701493
LMArena Hard Prompts14831516
LMArena Coding15081537
LMArena WebDev1652
LMArena Vision12601315
LMArena Agent4.2
AA Intelligence Index48.1
AA-LCR83.0%
AA-Omniscience25
GPQA Diamond (AA)93.5%
Humanity's Last Exam (AA)48.7%
SciCode (AA)58.8%
IOI56.6%
Terminal-Bench 2.1 (Vals)79.0%
Vals Index64.5
CritPt24.9%
GDPval (AA)58.7%
τ²-Bench Banking (AA)50.5%
Code Migration47.4%
Excel Modeling Benchmark67.4%
Finance Agent v260.0%
Harvey's Legal Agent Benchmark23.8%
Legal Research Bench55.3%
MysteryMechanism36.0%
Terminal-Bench 4.0 (Vals)27.8%
Terminal-Bench Science14.3%
Vibe Code Bench 1-10020.5%
Vibe Code Bench v1.185.9%
LMArena Maths14531496
LMArena Creative Writing14481453
LMArena Instruction Following14431483
LMArena Multi-turn14731493
LMArena Longer Queries14571506
LMArena Search1205
LMArena Document1468
Chess Puzzles38.0%
Mystery Game Puzzles25.0%

Data as of 2026-09-21. Best configuration of each model; every score links to its source on the model pages.

Grok 4.20 Multi-Agent vs Muse Spark 1.3: questions

Is Grok 4.20 Multi-Agent better than Muse Spark 1.3?
Muse Spark 1.3 (max) leads on quality: 69.4 vs 59.2. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 (max) is ahead overall as of 2026-09-21, but check the category scores for your use.
Is Grok 4.20 Multi-Agent better than Muse Spark 1.3 for coding?
Muse Spark 1.3 scores higher in coding (67 vs 66 on the category index, where 50 is average).
Which is cheaper, Grok 4.20 Multi-Agent or Muse Spark 1.3?
Grok 4.20 Multi-Agent is cheaper: $1.56 against $2.00 per million tokens, blended at three input tokens per output token.
Which is faster, Grok 4.20 Multi-Agent or Muse Spark 1.3?
Grok 4.20 Multi-Agent streams faster: 304 against 241 output tokens per second.
Which has the larger context window?
Muse Spark 1.3 accepts more context: 1.0M against 1M tokens.