BenchLeader

GPT-5 vs Muse Spark 1.3

Verdict
  • Muse Spark 1.3 (xhigh) leads on quality: 63.3 vs 60.8.
  • GPT-5 is stronger in coding, human preference, instruction following, maths, reasoning.
  • Muse Spark 1.3 (xhigh) is stronger in agents & tools, composite, knowledge, long context, multimodal.
  • Muse Spark 1.3 (xhigh) is 1.7× cheaper ($2.00 vs $3.44 per 1M blended).
  • Muse Spark 1.3 (xhigh) streams 2.4× faster (206 vs 85 tokens per second).
MetricGPT-5Muse Spark 1.3 (xhigh)
BenchLeader Index60.863.3
Agents & tools score51.060.2
Coding score60.960.5
Composite score58.178.7
Human preference score61.7
Instruction following score65.0
Knowledge score62.275.3
Long context score65.668.1
Maths score72.8
Multimodal score59.666.6
Reasoning score64.2
Blended price $/M$3.44$2.00
Output speed85 tok/s206 tok/s
Time to first answer60.7 s31.1 s
Context window400k1.0M
Terminal-Bench49.6%
SciCode42.9%
Remote Labor Index1.7%
WeirdML39.8%
APEX-Agents18.3%
Epoch Capabilities Index150
LMArena Text1427
LMArena Hard Prompts1449
LMArena Coding1462
LMArena WebDev1623
LMArena Vision1232
LiveBench81.6%
LiveBench Reasoning89.7%
LiveBench Coding81.1%
LiveBench Agentic Coding64.1%
LiveBench Mathematics96.0%
LiveBench Data Analysis79.6%
LiveBench Language82.8%
AA Intelligence Index23.045.2
IFBench73.1%
AA-LCR78.2%83.0%
MMMU-Pro74.2%82.0%
AA-Omniscience-8.723.1
Terminal-Bench Hard32.6%
GPQA Diamond (AA)85.3%94.1%
Humanity's Last Exam (AA)28.5%47.5%
SciCode (AA)59.7%
τ²-Bench Telecom (AA)84.8%
IOI43.9%
Terminal-Bench 2.1 (Vals)72.3%
Vals Index60.3
PRBench Finance51.3%
PRBench Legal49.0%
VISTA49.7%
MultiNRC52.1%
TutorBench55.3%
Kagi LLM Benchmark72.7%
IFEval (HELM)87.5%
Omni-MATH (HELM)64.7%
WildBench (HELM)85.7%
MMLU-Pro (HELM)86.3%
GPQA Diamond (HELM)79.1%
HELM Capabilities mean80.7%
Aider Polyglot88.0%
SWE-bench Verified (any scaffold)75.6%

Data as of 2026-09-12. Best configuration of each model; every score links to its source on the model pages.

GPT-5 vs Muse Spark 1.3: questions

Is GPT-5 better than Muse Spark 1.3?
Muse Spark 1.3 (xhigh) leads on quality: 63.3 vs 60.8. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 (xhigh) is ahead overall as of 2026-09-12, but check the category scores for your use.
Is GPT-5 better than Muse Spark 1.3 for coding?
GPT-5 scores higher in coding (61 vs 61 on the category index, where 50 is average).
Is GPT-5 better than Muse Spark 1.3 for agentic tasks?
Muse Spark 1.3 scores higher in agentic tasks (60 vs 51 on the category index, where 50 is average).
Which is cheaper, GPT-5 or Muse Spark 1.3?
Muse Spark 1.3 is cheaper: $2.00 against $3.44 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5 or Muse Spark 1.3?
Muse Spark 1.3 streams faster: 206 against 85 output tokens per second.
Which has the larger context window?
Muse Spark 1.3 accepts more context: 1.0M against 400k tokens.