BenchLeader

Grok 4.3 vs Muse Spark 1.3

Verdict
  • Muse Spark 1.3 (xhigh) leads on quality: 63.9 vs 60.7.
  • Grok 4.3 (medium) is stronger in instruction following.
  • Muse Spark 1.3 (xhigh) is stronger in agents & tools, composite, knowledge, long context, multimodal, coding.
  • Grok 4.3 (medium) is 1.3× cheaper ($1.56 vs $2.00 per 1M blended).
  • Muse Spark 1.3 (xhigh) streams 1.7× faster (190 vs 112 tokens per second).
MetricGrok 4.3 (medium)Muse Spark 1.3 (xhigh)
BenchLeader Index60.763.9
Agents & tools score60.160.2
Composite score60.378.6
Instruction following score80.0
Knowledge score72.175.1
Long context score64.068.1
Multimodal score60.266.6
Coding score70.3
Blended price $/M$1.56$2.00
Output speed112 tok/s190 tok/s
Time to first answer12.1 s35.7 s
Context window1M1M
LMArena WebDev1625
LiveBench81.6%
LiveBench Reasoning89.7%
LiveBench Coding81.1%
LiveBench Agentic Coding64.1%
LiveBench Mathematics96.0%
LiveBench Data Analysis79.6%
LiveBench Language82.8%
AA Intelligence Index24.845.2
IFBench83.3%
AA-LCR75.0%83.0%
MMMU-Pro75.8%82.0%
AA-Omniscience16.723.1
Terminal-Bench Hard30.3%
GPQA Diamond (AA)89.0%94.1%
Humanity's Last Exam (AA)30.0%47.5%
SciCode (AA)59.7%
τ²-Bench Telecom (AA)91.2%
Terminal-Bench 2.1 (Vals)72.3%
Vals Index60.3

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Grok 4.3 vs Muse Spark 1.3: questions

Is Grok 4.3 better than Muse Spark 1.3?
Muse Spark 1.3 (xhigh) leads on quality: 63.9 vs 60.7. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 (xhigh) is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Grok 4.3 better than Muse Spark 1.3 for agentic tasks?
Muse Spark 1.3 scores higher in agentic tasks (60 vs 60 on the category index, where 50 is average).
Which is cheaper, Grok 4.3 or Muse Spark 1.3?
Grok 4.3 is cheaper: $1.56 against $2.00 per million tokens, blended at three input tokens per output token.
Which is faster, Grok 4.3 or Muse Spark 1.3?
Muse Spark 1.3 streams faster: 190 against 112 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.