BenchLeader

Grok 4 vs Muse Spark 1.3

Verdict
  • Muse Spark 1.3 (xhigh) leads on quality: 62.0 vs 59.8.
  • Grok 4 is stronger in coding, human preference, instruction following, long context, maths, reasoning.
  • Muse Spark 1.3 (xhigh) is stronger in agents & tools, knowledge, multimodal, composite.
  • Grok 4 is 1.3× cheaper ($1.56 vs $2.00 per 1M blended).
MetricGrok 4Muse Spark 1.3 (xhigh)
BenchLeader Index59.862.0
Agents & tools score53.460.2
Coding score61.260.5
Human preference score59.8
Instruction following score67.7
Knowledge score56.071.8
Long context score77.765.8
Maths score58.0
Multimodal score54.565.7
Reasoning score58.3
Composite score75.0
Blended price $/M$1.56$2.00
Output speed235 tok/s
Time to first answer28.5 s
Context window256k1.0M
GPQA Diamond87.0%
OTIS Mock AIME84.0%
Terminal-Bench27.2%
SimpleBench60.5%
Fiction.LiveBench 120k96.9%
Cybench43.0%
WeirdML45.7%
APEX-Agents15.2%
Epoch Capabilities Index146.4
LMArena Text1411
LMArena Hard Prompts1420
LMArena Coding1435
LMArena WebDev1623
LMArena Vision1210
LiveBench81.6%
LiveBench Reasoning89.7%
LiveBench Coding81.1%
LiveBench Agentic Coding64.1%
LiveBench Mathematics96.0%
LiveBench Data Analysis79.6%
LiveBench Language82.8%
AA Intelligence Index45.2
AA-LCR83.0%
MMMU-Pro82.0%
AA-Omniscience23.1
GPQA Diamond (AA)94.1%
Humanity's Last Exam (AA)47.5%
SciCode (AA)59.7%
AIME (Vals)90.6%
LiveCodeBench83.3%
MMLU-Pro85.3%
IOI43.9%
LegalBench83.2%
CorpFin66.0%
TaxEval65.1%
MedQA92.5%
MGSM90.9%
Terminal-Bench 2.1 (Vals)72.3%
SWE-bench (Vals)57.8%
GPQA Diamond (Vals)88.1%
Vals Index60.3
IMO 202521.4%
MathArena Apex2.1%
Kagi LLM Benchmark73.6%
IFEval (HELM)94.9%
Omni-MATH (HELM)60.3%
WildBench (HELM)79.7%
MMLU-Pro (HELM)85.1%
GPQA Diamond (HELM)72.6%
HELM Capabilities mean78.5%
Aider Polyglot79.6%
ARC-AGI-179.6%
ARC-AGI-229.4%
BFCL Overall63.0%

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

Grok 4 vs Muse Spark 1.3: questions

Is Grok 4 better than Muse Spark 1.3?
Muse Spark 1.3 (xhigh) leads on quality: 62.0 vs 59.8. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 (xhigh) is ahead overall as of 2026-09-13, but check the category scores for your use.
Is Grok 4 better than Muse Spark 1.3 for coding?
Grok 4 scores higher in coding (61 vs 61 on the category index, where 50 is average).
Is Grok 4 better than Muse Spark 1.3 for agentic tasks?
Muse Spark 1.3 scores higher in agentic tasks (60 vs 53 on the category index, where 50 is average).
Which is cheaper, Grok 4 or Muse Spark 1.3?
Grok 4 is cheaper: $1.56 against $2.00 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Muse Spark 1.3 accepts more context: 1.0M against 256k tokens.