BenchLeader

Grok 4.6 vs Muse Spark 1.2

Verdict
  • Grok 4.6 (medium) leads on quality: 63.8 vs 62.6.
  • Grok 4.6 (medium) is stronger in composite, knowledge, long context.
  • Muse Spark 1.2 is stronger in coding, reasoning, maths.
  • Muse Spark 1.2 is 1.5× cheaper ($2.00 vs $3.00 per 1M blended).
  • Muse Spark 1.2 streams 4.0× faster (219 vs 55 tokens per second).
MetricGrok 4.6 (medium)Muse Spark 1.2
BenchLeader Index63.862.6
Coding score63.966.5
Composite score83.279.1
Knowledge score77.376.9
Long context score67.166.0
Reasoning score62.170.2
Maths score54.3
Blended price $/M$3.00$2.00
Output speed55 tok/s219 tok/s
Time to first answer36.7 s23.7 s
Context window500k1.0M
SimpleBench74.5%
SciCode54.6%
ProofBench43.0%
Epoch Capabilities Index155.5
AA Intelligence Index43.039.8
AA-LCR81.0%79.0%
AA-Omniscience2827.2
GPQA Diamond (AA)93.5%90.4%
Humanity's Last Exam (AA)42.1%45.5%
SciCode (AA)55.9%57.4%
IOI49.5%
ARC-AGI-187.5%
ARC-AGI-261.3%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.