BenchLeader

GPT-6.1 Sol vs Muse Spark 1.3

Verdict
  • GPT-6.1 Sol (max) leads on quality: 69.3 vs 67.7.
  • GPT-6.1 Sol (max) is stronger in coding, knowledge, maths, multimodal.
  • Muse Spark 1.3 (max) is stronger in agents & tools, composite, human preference, reasoning.
  • Muse Spark 1.3 (max) is 2.0× cheaper ($2.00 vs $4.00 per 1M blended).
  • Muse Spark 1.3 (max) streams 3.2× faster (175 vs 56 tokens per second).
MetricGPT-6.1 Sol (max)Muse Spark 1.3 (max)
BenchLeader Index69.367.7
Agents & tools score63.970.3
Coding score72.464.6
Composite score79.084.1
Human preference score68.169.4
Knowledge score78.872.9
Long context score66.866.8
Maths score72.563.6
Multimodal score66.265.9
Reasoning score76.478.1
Blended price $/M$4.00$2.00
Output speed56 tok/s175 tok/s
Time to first answer326.9 s70.5 s
Context window1.1M1.0M
GPQA Diamond95.4%–
FrontierMath Tiers 1–393.7%74.0%
FrontierMath Tier 4100.0%46.3%
OTIS Mock AIME100.0%–
SimpleQA Verified73.9%–
Terminal-Bench58.2%–
SciCode54.2%58.8%
APEX-Agents60.0%–
ProofBench–58.0%
LMArena Text14841494
LMArena Hard Prompts15071517
LMArena Coding15451538
LMArena WebDev17551657
LMArena Vision12881309
LMArena Agent11.74
LiveBench81.6%–
LiveBench Reasoning92.6%–
LiveBench Coding80.4%–
LiveBench Agentic Coding54.5%–
LiveBench Mathematics96.8%–
LiveBench Data Analysis82.7%–
LiveBench Language90.1%–
LiveBench Instruction Following74.2%–
AA Intelligence Index v4.3.251.848.1
AA-LCR83.0%83.0%
MMMU-Pro86.0%–
AA-Omniscience41.525
GPQA Diamond (AA)–93.5%
Humanity's Last Exam (AA)52.9%48.7%
SciCode (AA)54.2%58.8%
IOI96.9%56.6%
Terminal-Bench 2.1 (Vals)–79.0%
Vals Index61.158.2
ARC-AGI-196.5%–
ARC-AGI-294.2%–
ARC-AGI-396.2%–
CritPt31.7%24.9%
GDPval-AA v2.153.8%59.0%
τ³-Banking (AA)–50.5%
ITBench SRE (AA)–33.2%
Analyst Agent (AA)50.0%–
BioMysteryBench79.6%–
Code Migration65.1%47.4%
CUA-bench–5.8%
CyberBench39.3%72.7%
Excel Modeling Benchmark70.8%67.4%
Finance Agent v252.0%60.0%
Harvey's Legal Agent Benchmark5.4%23.8%
Legal Research Bench38.5%55.3%
MedCode48.8%–
MedScribe86.5%–
MysteryMechanism46.4%36.0%
Public Benefits Bench59.3%–
SAGE46.5%–
SREBench50.8%–
Tax Agent Bench62.3%72.4%
Terminal-Bench 4.0 (Vals)55.0%24.8%
Terminal-Bench Science–10.0%
Vibe Code Bench 1-100–20.5%
Vibe Code Bench v1.188.9%85.9%
LMArena Maths14881502
LMArena Creative Writing14621456
LMArena Instruction Following14881483
LMArena Multi-turn14871487
LMArena Longer Queries14961497
LMArena Document–1468
Chess Puzzles61.0%38.0%
EBR-bench54.3%–
Mystery Game Puzzles80.0%25.0%
CursorBench–41.6%
Terminal-Bench 4.0 (AA)56.1%33.3%
Terminal-Bench 2.1 (AA)–84.3%
AutomationBench64.9%57.9%
GDP.pdf31.0%26.6%
MLCR33.9%43.3%
Harvey LAB6.9%8.9%
AA-Omniscience: accuracy62.1%43.6%
AA-Omniscience: non-hallucination45.7%67.1%
AA-Briefcase v1.115571581

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

GPT-6.1 Sol vs Muse Spark 1.3: questions

Is GPT-6.1 Sol better than Muse Spark 1.3?
GPT-6.1 Sol (max) leads on quality: 69.3 vs 67.7. The BenchLeader Index combines every independent quality benchmark; GPT-6.1 Sol (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is GPT-6.1 Sol better than Muse Spark 1.3 for coding?
GPT-6.1 Sol scores higher in coding (72 vs 65 on the category index, where 50 is average).
Is GPT-6.1 Sol better than Muse Spark 1.3 for agentic tasks?
Muse Spark 1.3 scores higher in agentic tasks (70 vs 64 on the category index, where 50 is average).
Which is cheaper, GPT-6.1 Sol or Muse Spark 1.3?
Muse Spark 1.3 is cheaper: $2.00 against $4.00 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-6.1 Sol or Muse Spark 1.3?
Muse Spark 1.3 streams faster: 175 against 56 output tokens per second.
Which has the larger context window?
GPT-6.1 Sol accepts more context: 1.1M against 1.0M tokens.