BenchLeader

Grok 4.7 vs Muse Spark 1.1

Verdict
  • Grok 4.7 (high) and Muse Spark 1.1 are level on quality (65.8 vs 65.2).
  • Grok 4.7 (high) is stronger in composite, knowledge, long context, reasoning.
  • Muse Spark 1.1 is stronger in agents & tools, coding, human preference, instruction following, maths, multimodal.
  • Muse Spark 1.1 is 1.5× cheaper ($2.00 vs $3.00 per 1M blended).
  • Muse Spark 1.1 streams 3.8× faster (210 vs 55 tokens per second).
MetricGrok 4.7 (high)Muse Spark 1.1
BenchLeader Index65.865.2
Composite score87.5
Knowledge score72.670.5
Long context score64.8
Reasoning score76.469.3
Agents & tools score62.6
Coding score68.0
Human preference score66.2
Instruction following score77.8
Maths score62.7
Multimodal score64.8
Blended price $/M$3.00$2.00
Output speed55 tok/s210 tok/s
Time to first answer1.4 s2.8 s
Context window500k1.0M
SimpleQA Verified57.8%
SciCode58.2%
APEX-Agents41.9%
ProofBench39.0%
Epoch Capabilities Index154.6
LMArena Text1493
LMArena Hard Prompts1513
LMArena Coding1531
LMArena WebDev1542
LMArena Vision1294
LMArena Agent-3
AA Intelligence Index46.3
AA-LCR77.0%
AA-Omniscience30.9
Humanity's Last Exam (AA)42.3%
SciCode (AA)57.8%
LegalBench85.1%
SWE-Bench Pro61.5%
MCP Atlas88.1%
MultiChallenge75.3%
PRBench Finance55.0%
PRBench Legal57.0%
MultiNRC65.6%
EQ-Bench 41260
CritPt18.0%
GDPval (AA)59.7%
Code Migration34.7%
Excel Modeling Benchmark54.9%
Harvey's Legal Agent Benchmark19.6%
Legal Research Bench39.9%
MedCode48.7%
MedScribe87.2%
ProgramBench0.0%
Public Benefits Bench68.5%
SAGE31.0%
FORTRESS12.4%
SWE-Bench Pro (private)51.5%
VTB44.8%
LMArena Maths1486
LMArena Creative Writing1451
LMArena Instruction Following1471
LMArena Multi-turn1497
LMArena Longer Queries1481
LMArena Document1480
DeepSWE53.3%
GBAEval7.9%
Vending-Bench 26520.5
GDP.pdf15.0%

Data as of 2026-09-21. Best configuration of each model; every score links to its source on the model pages.

Grok 4.7 vs Muse Spark 1.1: questions

Is Grok 4.7 better than Muse Spark 1.1?
Grok 4.7 (high) and Muse Spark 1.1 are level on quality (65.8 vs 65.2). The BenchLeader Index combines every independent quality benchmark; Grok 4.7 (high) is ahead overall as of 2026-09-21, but check the category scores for your use.
Which is cheaper, Grok 4.7 or Muse Spark 1.1?
Muse Spark 1.1 is cheaper: $2.00 against $3.00 per million tokens, blended at three input tokens per output token.
Which is faster, Grok 4.7 or Muse Spark 1.1?
Muse Spark 1.1 streams faster: 210 against 55 output tokens per second.
Which has the larger context window?
Muse Spark 1.1 accepts more context: 1.0M against 500k tokens.