BenchLeader

Grok 4.5 vs Muse Spark

Verdict
  • Muse Spark leads on quality: 65.4 vs 63.4.
  • Grok 4.5 is stronger in knowledge, long context.
  • Muse Spark is stronger in agents & tools, coding, composite, human preference, multimodal, reasoning, instruction following, maths.
MetricGrok 4.5Muse Spark
BenchLeader Index63.465.4
Agents & tools score58.666.3
Coding score59.966.0
Composite score66.068.4
Human preference score66.969.0
Knowledge score76.064.5
Long context score66.265.5
Multimodal score64.865.9
Reasoning score68.970.0
Instruction following score79.6
Maths score56.3
Blended price $/M$3.00
Output speed56 tok/s
Time to first answer12.7 s
Context window500k262k
GPQA Diamond89.8%
OTIS Mock AIME88.9%
Humanity's Last Exam40.6%
SimpleBench70.0%
SciCode51.5%
WeirdML46.4%
APEX-Agents34.2%
FrontierCode42.4%
ProofBench17.0%
Epoch Capabilities Index153.9152.1
LMArena Text14711488
LMArena Hard Prompts14951505
LMArena Coding15231526
LMArena WebDev1556
LMArena Vision12911306
LMArena Agent3.9
LiveBench75.8%
LiveBench Reasoning87.2%
LiveBench Coding68.6%
LiveBench Agentic Coding56.5%
LiveBench Mathematics90.8%
LiveBench Data Analysis73.0%
LiveBench Language82.8%
AA Intelligence Index39.131.3
IFBench75.9%
AA-LCR79.3%78.0%
MMMU-Pro80.4%80.5%
AA-Omniscience25.37.2
Terminal-Bench Hard45.5%
GPQA Diamond (AA)93.1%88.4%
Humanity's Last Exam (AA)42.7%40.7%
SciCode (AA)55.0%
τ²-Bench Telecom (AA)91.5%
AIME (Vals)96.9%
MMLU-Pro87.3%
LegalBench84.2%
CorpFin65.1%
TaxEval77.7%
SWE-bench (Vals)74.4%
GPQA Diamond (Vals)89.7%
SWE-Bench Pro55.0%
MCP Atlas82.2%
MultiChallenge75.5%
PRBench Finance52.4%
PRBench Legal52.3%
MultiNRC59.0%
TutorBench68.5%
Kagi LLM Benchmark83.5%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.