BenchLeader

GPT-5.5 Instant vs Muse Spark

Verdict
  • Muse Spark leads on quality: 65.6 vs 60.5.
  • GPT-5.5 Instant is stronger in agents & tools, knowledge.
  • Muse Spark is stronger in coding, composite, human preference, instruction following, long context, maths, multimodal, reasoning.
MetricGPT-5.5 InstantMuse Spark
BenchLeader Index60.565.6
Agents & tools score70.766.4
Coding score61.166.2
Composite score62.968.6
Human preference score67.569.3
Instruction following score69.879.7
Knowledge score69.464.6
Long context score61.465.5
Maths score42.556.4
Multimodal score59.765.9
Reasoning score62.670.2
Blended price $/M$11.25
Output speed129 tok/s
Time to first answer16.8 s
Context window400k262k
GPQA Diamond82.5%89.8%
FrontierMath Tiers 1–326.3%
FrontierMath Tier 42.4%
OTIS Mock AIME68.1%88.9%
Humanity's Last Exam40.6%
SciCode48.6%51.5%
ProofBench17.0%
Epoch Capabilities Index142.5152.1
LMArena Text14741488
LMArena Hard Prompts14911505
LMArena Coding15141526
LMArena Vision12511306
AA Intelligence Index26.831.3
IFBench71.5%75.9%
AA-LCR70.0%78.0%
MMMU-Pro80.5%
AA-Omniscience11.27.2
Terminal-Bench Hard42.4%45.5%
GPQA Diamond (AA)84.7%88.4%
Humanity's Last Exam (AA)21.6%40.7%
SciCode (AA)52.5%
τ²-Bench Telecom (AA)49.4%91.5%
AIME (Vals)96.9%
MMLU-Pro87.3%
LegalBench84.2%
CorpFin65.1%
TaxEval77.7%
SWE-bench (Vals)74.4%
GPQA Diamond (Vals)89.7%
SWE-Bench Pro55.0%
MCP Atlas82.2%
MultiChallenge75.5%
PRBench Finance52.4%
PRBench Legal52.3%
MultiNRC59.0%
TutorBench68.5%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

GPT-5.5 Instant vs Muse Spark: questions

Is GPT-5.5 Instant better than Muse Spark?
Muse Spark leads on quality: 65.6 vs 60.5. The BenchLeader Index combines every independent quality benchmark; Muse Spark is ahead overall as of 2026-09-10, but check the category scores for your use.
Is GPT-5.5 Instant better than Muse Spark for coding?
Muse Spark scores higher in coding (66 vs 61 on the category index, where 50 is average).
Is GPT-5.5 Instant better than Muse Spark for agentic tasks?
GPT-5.5 Instant scores higher in agentic tasks (71 vs 66 on the category index, where 50 is average).
Which has the larger context window?
GPT-5.5 Instant accepts more context: 400k against 262k tokens.