BenchLeader

GPT-5-Codex vs Muse Spark 1.1

Verdict
  • Muse Spark 1.1 leads on quality: 65.0 vs 58.2.
  • Muse Spark 1.1 is stronger in agents & tools, coding, composite, instruction following, knowledge, long context, multimodal, reasoning, human preference, maths.
  • Muse Spark 1.1 is 1.7× cheaper ($2.00 vs $3.44 per 1M blended).
MetricGPT-5-CodexMuse Spark 1.1
BenchLeader Index58.265.0
Agents & tools score56.362.7
Coding score53.668.1
Composite score60.472.3
Instruction following score72.177.8
Knowledge score60.673.2
Long context score62.265.4
Multimodal score58.265.0
Reasoning score64.369.2
Human preference score66.3
Maths score52.2
Blended price $/M$3.44$2.00
Output speed194 tok/s
Time to first answer2.6 s
Context window400k1.0M
SimpleQA Verified57.8%
Terminal-Bench44.3%
SciCode58.2%
WeirdML54.5%
APEX-Agents41.9%
ProofBench39.0%
Epoch Capabilities Index154.6
LMArena Text1492
LMArena Hard Prompts1511
LMArena Coding1531
LMArena WebDev1541
LMArena Vision1293
LMArena Agent-2.6
AA Intelligence Index24.934.3
IFBench74.2%
AA-LCR71.7%77.7%
MMMU-Pro73.8%
AA-Omniscience-7.528.1
Terminal-Bench Hard37.9%
GPQA Diamond (AA)83.7%89.8%
Humanity's Last Exam (AA)27.9%46.2%
SciCode (AA)58.8%
τ²-Bench Telecom (AA)86.8%
SWE-Bench Pro61.5%
MCP Atlas88.1%
MultiChallenge75.3%
PRBench Finance55.0%
PRBench Legal57.0%
MultiNRC65.6%
EQ-Bench 41260
Kagi LLM Benchmark70.3%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

GPT-5-Codex vs Muse Spark 1.1: questions

Is GPT-5-Codex better than Muse Spark 1.1?
Muse Spark 1.1 leads on quality: 65.0 vs 58.2. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.1 is ahead overall as of 2026-09-10, but check the category scores for your use.
Is GPT-5-Codex better than Muse Spark 1.1 for coding?
Muse Spark 1.1 scores higher in coding (68 vs 54 on the category index, where 50 is average).
Is GPT-5-Codex better than Muse Spark 1.1 for agentic tasks?
Muse Spark 1.1 scores higher in agentic tasks (63 vs 56 on the category index, where 50 is average).
Which is cheaper, GPT-5-Codex or Muse Spark 1.1?
Muse Spark 1.1 is cheaper: $2.00 against $3.44 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Muse Spark 1.1 accepts more context: 1.0M against 400k tokens.