BenchLeader

GPT-6 Astra vs Muse Spark 1.3

Verdict
  • GPT-6 Astra (max) leads on quality: 70.6 vs 63.9.
  • GPT-6 Astra (max) is stronger in agents & tools, coding, knowledge, maths, reasoning.
  • Muse Spark 1.3 (xhigh) is stronger in composite, long context, multimodal.
  • Muse Spark 1.3 (xhigh) is 11× cheaper ($2.00 vs $22.00 per 1M blended).
  • Muse Spark 1.3 (xhigh) streams 3.6× faster (193 vs 54 tokens per second).
MetricGPT-6 Astra (max)Muse Spark 1.3 (xhigh)
BenchLeader Index70.663.9
Agents & tools score68.260.4
Coding score79.070.8
Composite score73.178.6
Knowledge score79.875.0
Maths score77.0
Reasoning score70.9
Long context score68.1
Multimodal score66.6
Blended price $/M$22.00$2.00
Output speed54 tok/s193 tok/s
Time to first answer328.6 s40.0 s
Context window1.1M1M
GPQA Diamond95.8%
FrontierMath Tiers 1–393.7%
FrontierMath Tier 497.6%
OTIS Mock AIME100.0%
SimpleQA Verified75.6%
Terminal-Bench58.2%
SciCode56.5%
FrontierCode53.3%
LMArena WebDev17961625
LMArena Agent12.5
LiveBench82.2%81.6%
LiveBench Reasoning92.7%89.7%
LiveBench Coding80.4%81.1%
LiveBench Agentic Coding57.3%64.1%
LiveBench Mathematics96.8%96.0%
LiveBench Data Analysis83.0%79.6%
LiveBench Language89.4%82.8%
AA Intelligence Index45.2
AA-LCR83.0%
MMMU-Pro82.0%
AA-Omniscience23.1
GPQA Diamond (AA)94.1%
Humanity's Last Exam (AA)47.5%
SciCode (AA)59.7%
Terminal-Bench 2.1 (Vals)87.3%72.3%
Vals Index66.660.3
ARC-AGI-197.5%
ARC-AGI-295.0%
ARC-AGI-362.7%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.