BenchLeader

GPT-6.1 Sol vs Muse Spark

Verdict
  • GPT-6.1 Sol (max) leads on quality: 69.3 vs 64.6.
  • GPT-6.1 Sol (max) is stronger in coding, composite, knowledge, long context, maths, multimodal, reasoning.
  • Muse Spark is stronger in agents & tools, human preference, instruction following.
MetricGPT-6.1 Sol (max)Muse Spark
BenchLeader Index69.364.6
Agents & tools score63.966.0
Coding score72.465.3
Composite score79.065.0
Human preference score68.168.7
Knowledge score78.863.7
Long context score66.864.2
Maths score72.559.3
Multimodal score66.264.6
Reasoning score76.465.9
Instruction following score–78.3
Blended price $/M$4.00–
Output speed56 tok/s–
Time to first answer326.9 s–
Context window1.1M262k
GPQA Diamond95.4%89.8%
FrontierMath Tiers 1–393.7%–
FrontierMath Tier 4100.0%–
OTIS Mock AIME100.0%88.9%
SimpleQA Verified73.9%–
Humanity's Last Exam–40.6%
Terminal-Bench58.2%–
SciCode54.2%51.5%
APEX-Agents60.0%–
ProofBench–17.0%
Epoch Capabilities Index–152.0
LMArena Text14841489
LMArena Hard Prompts15071506
LMArena Coding15451529
LMArena WebDev1755–
LMArena Vision12881306
LMArena Agent11.7–
LiveBench81.6%–
LiveBench Reasoning92.6%–
LiveBench Coding80.4%–
LiveBench Agentic Coding54.5%–
LiveBench Mathematics96.8%–
LiveBench Data Analysis82.7%–
LiveBench Language90.1%–
LiveBench Instruction Following74.2%–
AA Intelligence Index v4.3.251.831.3
IFBench–75.9%
AA-LCR83.0%78.0%
MMMU-Pro86.0%80.5%
AA-Omniscience41.57.2
Terminal-Bench Hard–45.5%
GPQA Diamond (AA)–88.4%
Humanity's Last Exam (AA)52.9%40.7%
SciCode (AA)54.2%–
τ²-Bench Telecom (AA)–91.5%
AIME (Vals)–96.9%
MMLU-Pro–87.3%
IOI96.9%–
LegalBench–84.2%
CorpFin–65.1%
TaxEval–77.7%
SWE-bench (Vals)–74.4%
GPQA Diamond (Vals)–89.7%
Vals Index61.1–
SWE-Bench Pro–55.0%
MCP Atlas–82.2%
MultiChallenge–75.5%
PRBench Finance–52.4%
PRBench Legal–52.3%
MultiNRC–59.0%
TutorBench–68.5%
ARC-AGI-196.5%–
ARC-AGI-294.2%–
ARC-AGI-396.2%–
CritPt31.7%11.3%
GDPval-AA v2.153.8%25.1%
Analyst Agent (AA)50.0%–
BioMysteryBench79.6%–
CaseLaw v2–63.1%
Code Migration65.1%–
CyberBench39.3%–
Excel Modeling Benchmark70.8%–
Finance Agent v252.0%–
Harvey's Legal Agent Benchmark5.4%–
Legal Research Bench38.5%–
MedCode48.8%51.3%
MedScribe86.5%85.9%
MMMU-Pro (Vals)–87.4%
MysteryMechanism46.4%–
Public Benefits Bench59.3%–
SAGE46.5%–
SREBench50.8%–
Tax Agent Bench62.3%–
Terminal-Bench 2.0 (Vals)–59.5%
Terminal-Bench 4.0 (Vals)55.0%–
Vibe Code Bench v1.188.9%19.7%
FORTRESS–20.2%
SWE-Bench Pro (private)–55.0%
SWE Atlas: Codebase QnA–24.2%
SWE Atlas: Test Writing–31.1%
LMArena Maths14881466
LMArena Creative Writing14621465
LMArena Instruction Following14881464
LMArena Multi-turn14871493
LMArena Longer Queries14961476
LMArena Document–1467
Chess Puzzles61.0%–
EBR-bench54.3%–
Mystery Game Puzzles80.0%–
Terminal-Bench 4.0 (AA)56.1%–
Terminal-Bench 2.1 (AA)–62.2%
AutomationBench64.9%–
GDP.pdf31.0%–
MLCR33.9%–
Harvey LAB6.9%–
AA-Omniscience: accuracy62.1%49.6%
AA-Omniscience: non-hallucination45.7%15.8%
AA-Briefcase v1.11557–

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

GPT-6.1 Sol vs Muse Spark: questions

Is GPT-6.1 Sol better than Muse Spark?
GPT-6.1 Sol (max) leads on quality: 69.3 vs 64.6. The BenchLeader Index combines every independent quality benchmark; GPT-6.1 Sol (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is GPT-6.1 Sol better than Muse Spark for coding?
GPT-6.1 Sol scores higher in coding (72 vs 65 on the category index, where 50 is average).
Is GPT-6.1 Sol better than Muse Spark for agentic tasks?
Muse Spark scores higher in agentic tasks (66 vs 64 on the category index, where 50 is average).
Which has the larger context window?
GPT-6.1 Sol accepts more context: 1.1M against 262k tokens.