BenchLeader

GPT-5.4 Pro vs Muse Spark 1.3

Verdict
  • GPT-5.4 Pro and Muse Spark 1.3 (xhigh) are level on quality (64.2 vs 63.9).
  • GPT-5.4 Pro is stronger in instruction following, multimodal, reasoning.
  • Muse Spark 1.3 (xhigh) is stronger in knowledge, agents & tools, coding, composite, long context.
  • Muse Spark 1.3 (xhigh) is 34× cheaper ($2.00 vs $67.50 per 1M blended).
  • Muse Spark 1.3 (xhigh) streams 193.4× faster (193 vs 1 tokens per second).
MetricGPT-5.4 ProMuse Spark 1.3 (xhigh)
BenchLeader Index64.263.9
Instruction following score67.3
Knowledge score73.275.0
Multimodal score69.866.6
Reasoning score75.1
Agents & tools score60.4
Coding score70.8
Composite score78.6
Long context score68.1
Blended price $/M$67.50$2.00
Output speed1 tok/s193 tok/s
Time to first answer6.7 s40.0 s
Context window1.1M1M
Humanity's Last Exam44.3%
SimpleBench74.1%
Epoch Capabilities Index158.9
LMArena WebDev1625
LiveBench81.6%
LiveBench Reasoning89.7%
LiveBench Coding81.1%
LiveBench Agentic Coding64.1%
LiveBench Mathematics96.0%
LiveBench Data Analysis79.6%
LiveBench Language82.8%
AA Intelligence Index45.2
AA-LCR83.0%
MMMU-Pro82.0%
AA-Omniscience23.1
GPQA Diamond (AA)94.1%
Humanity's Last Exam (AA)47.5%
SciCode (AA)59.7%
Terminal-Bench 2.1 (Vals)72.3%
Vals Index60.3
MultiChallenge69.2%
VISTA53.9%
MultiNRC62.3%
TutorBench56.6%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.