BenchLeader

GPT-5-Codex vs Muse Spark 1.3

Verdict
  • Muse Spark 1.3 leads on quality: 68.3 vs 58.4.
  • GPT-5-Codex is stronger in agents & tools, coding, instruction following, multimodal, reasoning.
  • Muse Spark 1.3 is stronger in composite, knowledge, long context.
  • Muse Spark 1.3 is 1.7× cheaper ($2.00 vs $3.44 per 1M blended).
MetricGPT-5-CodexMuse Spark 1.3
BenchLeader Index58.468.3
Agents & tools score56.4
Coding score53.6
Composite score60.790.3
Instruction following score72.6
Knowledge score60.879.1
Long context score62.468.2
Multimodal score58.4
Reasoning score64.3
Blended price $/M$3.44$2.00
Output speed241 tok/s
Time to first answer30.5 s
Context window400k1.0M
Terminal-Bench44.3%
WeirdML54.5%
AA Intelligence Index24.948.2
IFBench74.2%
AA-LCR71.7%83.0%
MMMU-Pro73.8%
AA-Omniscience-7.525
Terminal-Bench Hard37.9%
GPQA Diamond (AA)83.7%93.5%
Humanity's Last Exam (AA)27.9%48.7%
SciCode (AA)58.8%
τ²-Bench Telecom (AA)86.8%
PRBench Finance59.5%
PRBench Legal61.6%
Kagi LLM Benchmark70.3%

Data as of 2026-09-14. Best configuration of each model; every score links to its source on the model pages.

GPT-5-Codex vs Muse Spark 1.3: questions

Is GPT-5-Codex better than Muse Spark 1.3?
Muse Spark 1.3 leads on quality: 68.3 vs 58.4. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 is ahead overall as of 2026-09-14, but check the category scores for your use.
Which is cheaper, GPT-5-Codex or Muse Spark 1.3?
Muse Spark 1.3 is cheaper: $2.00 against $3.44 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Muse Spark 1.3 accepts more context: 1.0M against 400k tokens.