BenchLeader

GPT-5.2 vs Muse Spark 1.3

Verdict
  • Muse Spark 1.3 (xhigh) leads on quality: 63.9 vs 61.6.
  • GPT-5.2 is stronger in agents & tools, human preference, instruction following, reasoning.
  • Muse Spark 1.3 (xhigh) is stronger in coding, composite, knowledge, long context, multimodal.
  • Muse Spark 1.3 (xhigh) is 2.4× cheaper ($2.00 vs $4.81 per 1M blended).
  • Muse Spark 1.3 (xhigh) streams 2.8× faster (190 vs 68 tokens per second).
MetricGPT-5.2Muse Spark 1.3 (xhigh)
BenchLeader Index61.663.9
Agents & tools score61.260.2
Coding score53.770.3
Composite score67.578.6
Human preference score67.9
Instruction following score67.4
Knowledge score61.175.1
Long context score67.968.1
Multimodal score59.966.6
Reasoning score61.3
Blended price $/M$4.81$2.00
Output speed68 tok/s190 tok/s
Time to first answer129.4 s35.7 s
Context window400k1M
Humanity's Last Exam27.8%
Terminal-Bench64.9%
SimpleBench45.8%
Remote Labor Index2.1%
APEX-Agents23.0%
Epoch Capabilities Index153.5
LMArena Text1476
LMArena Hard Prompts1497
LMArena Coding1515
LMArena WebDev14171625
LMArena Vision1268
LiveBench81.6%
LiveBench Reasoning89.7%
LiveBench Coding81.1%
LiveBench Agentic Coding64.1%
LiveBench Mathematics96.0%
LiveBench Data Analysis79.6%
LiveBench Language82.8%
AA Intelligence Index30.445.2
IFBench75.4%
AA-LCR82.7%83.0%
MMMU-Pro82.0%
AA-Omniscience-0.923.1
Terminal-Bench Hard47.0%
GPQA Diamond (AA)90.3%94.1%
Humanity's Last Exam (AA)37.7%47.5%
SciCode (AA)59.7%
τ²-Bench Telecom (AA)84.8%
Terminal-Bench 2.1 (Vals)72.3%
Vals Index60.3
SWE-Bench Pro29.9%
VISTA46.6%
MultiNRC42.2%
TutorBench53.5%
Kagi LLM Benchmark73.3%
SWE-bench Verified (bash only)69.0%
ARC-AGI-194.5%
ARC-AGI-272.9%
BFCL Overall55.9%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

GPT-5.2 vs Muse Spark 1.3: questions

Is GPT-5.2 better than Muse Spark 1.3?
Muse Spark 1.3 (xhigh) leads on quality: 63.9 vs 61.6. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 (xhigh) is ahead overall as of 2026-09-10, but check the category scores for your use.
Is GPT-5.2 better than Muse Spark 1.3 for coding?
Muse Spark 1.3 scores higher in coding (70 vs 54 on the category index, where 50 is average).
Is GPT-5.2 better than Muse Spark 1.3 for agentic tasks?
GPT-5.2 scores higher in agentic tasks (61 vs 60 on the category index, where 50 is average).
Which is cheaper, GPT-5.2 or Muse Spark 1.3?
Muse Spark 1.3 is cheaper: $2.00 against $4.81 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.2 or Muse Spark 1.3?
Muse Spark 1.3 streams faster: 190 against 68 output tokens per second.
Which has the larger context window?
Muse Spark 1.3 accepts more context: 1M against 400k tokens.