BenchLeader

GPT-5.6 Luna vs Muse Spark 1.3

Verdict
  • Muse Spark 1.3 leads on quality: 68.3 vs 59.3.
  • GPT-5.6 Luna (xhigh) is stronger in agents & tools, coding, human preference, multimodal, reasoning.
  • Muse Spark 1.3 is stronger in composite, knowledge, long context.
  • GPT-5.6 Luna (xhigh) is 4.4× cheaper ($0.450 vs $2.00 per 1M blended).
  • Muse Spark 1.3 streams 2.1× faster (241 vs 118 tokens per second).
MetricGPT-5.6 Luna (xhigh)Muse Spark 1.3
BenchLeader Index59.368.3
Agents & tools score48.7
Coding score59.8
Composite score73.390.3
Human preference score64.8
Knowledge score59.379.1
Long context score67.668.2
Multimodal score61.8
Reasoning score53.9
Blended price $/M$0.450$2.00
Output speed118 tok/s241 tok/s
Time to first answer47.7 s30.5 s
Context window1.1M1.0M
SimpleBench46.8%
SciCode50.0%
LMArena Text1453
LMArena Hard Prompts1473
LMArena Coding1498
LMArena WebDev1519
LMArena Vision1259
LMArena Agent0.7
AA Intelligence Index34.848.2
AA-LCR81.7%83.0%
MMMU-Pro78.5%
AA-Omniscience-10.825
GPQA Diamond (AA)89.5%93.5%
Humanity's Last Exam (AA)37.0%48.7%
SciCode (AA)50.5%58.8%
PRBench Finance59.5%
PRBench Legal61.6%
ARC-AGI-187.7%
ARC-AGI-247.6%
ARC-AGI-30.0%

Data as of 2026-09-14. Best configuration of each model; every score links to its source on the model pages.

GPT-5.6 Luna vs Muse Spark 1.3: questions

Is GPT-5.6 Luna better than Muse Spark 1.3?
Muse Spark 1.3 leads on quality: 68.3 vs 59.3. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 is ahead overall as of 2026-09-14, but check the category scores for your use.
Which is cheaper, GPT-5.6 Luna or Muse Spark 1.3?
GPT-5.6 Luna is cheaper: $0.450 against $2.00 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.6 Luna or Muse Spark 1.3?
Muse Spark 1.3 streams faster: 241 against 118 output tokens per second.
Which has the larger context window?
GPT-5.6 Luna accepts more context: 1.1M against 1.0M tokens.