BenchLeader

Grok 4.20 vs Muse Spark 1.3

Verdict
  • Muse Spark 1.3 (max) leads on quality: 69.3 vs 59.9.
  • Grok 4.20 (thinking) is stronger in instruction following.
  • Muse Spark 1.3 (max) is stronger in agents & tools, coding, composite, human preference, knowledge, long context, maths, multimodal, reasoning.
  • Grok 4.20 (thinking) is 1.3× cheaper ($1.56 vs $2.00 per 1M blended).
  • Muse Spark 1.3 (max) streams 2.2× faster (223 vs 101 tokens per second).
MetricGrok 4.20 (thinking)Muse Spark 1.3 (max)
BenchLeader Index59.969.3
Agents & tools score51.771.1
Coding score54.766.6
Composite score61.590.1
Human preference score67.169.8
Instruction following score79.9
Knowledge score56.876.1
Long context score60.768.0
Maths score55.669.2
Multimodal score59.967.4
Reasoning score60.783.1
Blended price $/M$1.56$2.00
Output speed101 tok/s223 tok/s
Time to first answer21.4 s34.8 s
Context window1M1.0M
GPQA Diamond89.3%
FrontierMath Tiers 1–344.9%
FrontierMath Tier 417.1%
OTIS Mock AIME92.2%
SimpleQA Verified30.2%
ProofBench14.0%
LMArena Text14721493
LMArena Hard Prompts14881516
LMArena Coding15101537
LMArena WebDev13741652
LMArena Vision12631315
LMArena Agent4.2
AA Intelligence Index25.748.2
IFBench82.9%
AA-LCR69.0%83.0%
MMMU-Pro74.6%
AA-Omniscience14.825
Terminal-Bench Hard40.9%
GPQA Diamond (AA)91.1%93.5%
Humanity's Last Exam (AA)34.5%48.7%
SciCode (AA)58.8%
τ²-Bench Telecom (AA)96.5%
AIME (Vals)96.5%
LiveCodeBench84.3%
MMLU-Pro86.3%
IOI56.6%
LegalBench77.7%
CorpFin63.7%
TaxEval74.1%
MedQA94.5%
Terminal-Bench 2.1 (Vals)44.2%79.0%
SWE-bench (Vals)72.2%
GPQA Diamond (Vals)88.6%
Vals Index17.664.5
Kagi LLM Benchmark75.0%
ARC-AGI-189.5%
ARC-AGI-265.1%
ARC-AGI-30.1%
CritPt6.6%24.9%
GDPval (AA)60.2%
τ²-Bench Banking (AA)50.5%
APEX-Agents (AA)14.2%
CaseLaw v254.5%
Code Migration0.3%47.4%
Excel Modeling Benchmark11.7%67.4%
Finance Agent v228.5%60.0%
Harvey's Legal Agent Benchmark0.0%23.8%
Legal Research Bench13.9%55.3%
MedCode32.2%
MedScribe63.4%
MMMU-Pro (Vals)83.5%
MortgageTax45.4%
MysteryMechanism36.0%
SAGE38.2%
Terminal-Bench 2.0 (Vals)40.5%
Terminal-Bench 4.0 (Vals)27.8%
Terminal-Bench Science14.3%
Vals Multimodal Index39.1%
Vibe Code Bench 1-10020.5%
Vibe Code Bench v1.14.1%85.9%
LMArena Maths14671496
LMArena Creative Writing14461453
LMArena Instruction Following14451483
LMArena Multi-turn14801493
LMArena Longer Queries14651506
LMArena Document14391468
Chess Puzzles24.0%
ForecastBench60.7%

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

Grok 4.20 vs Muse Spark 1.3: questions

Is Grok 4.20 better than Muse Spark 1.3?
Muse Spark 1.3 (max) leads on quality: 69.3 vs 59.9. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 (max) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is Grok 4.20 better than Muse Spark 1.3 for coding?
Muse Spark 1.3 scores higher in coding (67 vs 55 on the category index, where 50 is average).
Is Grok 4.20 better than Muse Spark 1.3 for agentic tasks?
Muse Spark 1.3 scores higher in agentic tasks (71 vs 52 on the category index, where 50 is average).
Which is cheaper, Grok 4.20 or Muse Spark 1.3?
Grok 4.20 is cheaper: $1.56 against $2.00 per million tokens, blended at three input tokens per output token.
Which is faster, Grok 4.20 or Muse Spark 1.3?
Muse Spark 1.3 streams faster: 223 against 101 output tokens per second.
Which has the larger context window?
Muse Spark 1.3 accepts more context: 1.0M against 1M tokens.