BenchLeader

MiMo-V2.6-Pro vs Muse Spark

Verdict
  • Muse Spark leads on quality: 65.3 vs 64.3.
  • MiMo-V2.6-Pro is stronger in composite, knowledge, long context, reasoning.
  • Muse Spark is stronger in agents & tools, coding, human preference, instruction following, maths, multimodal.
MetricMiMo-V2.6-ProMuse Spark
BenchLeader Index64.365.3
Agents & tools score55.166.0
Composite score84.866.7
Knowledge score66.863.9
Long context score69.164.8
Reasoning score88.866.7
Coding score65.9
Human preference score69.2
Instruction following score79.8
Maths score60.0
Multimodal score65.3
Blended price $/M$0.548
Output speed54 tok/s
Time to first answer39.8 s
Context window1.0M262k
GPQA Diamond89.8%
OTIS Mock AIME88.9%
Humanity's Last Exam40.6%
SciCode51.5%
ProofBench17.0%
Epoch Capabilities Index152.1
LMArena Text1488
LMArena Hard Prompts1505
LMArena Coding1526
LMArena Vision1306
AA Intelligence Index46.331.3
IFBench75.9%
AA-LCR86.3%78.0%
MMMU-Pro80.5%
AA-Omniscience8.47.2
Terminal-Bench Hard45.5%
GPQA Diamond (AA)88.4%
Humanity's Last Exam (AA)49.4%40.7%
SciCode (AA)60.9%
τ²-Bench Telecom (AA)91.5%
AIME (Vals)96.9%
MMLU-Pro87.3%
LegalBench84.2%
CorpFin65.1%
TaxEval77.7%
Terminal-Bench 2.1 (Vals)67.8%
SWE-bench (Vals)74.4%
GPQA Diamond (Vals)89.7%
Vals Index59.7
SWE-Bench Pro55.0%
MCP Atlas82.2%
MultiChallenge75.5%
PRBench Finance52.4%
PRBench Legal52.3%
MultiNRC59.0%
TutorBench68.5%
CritPt26.6%11.3%
GDPval (AA)58.7%24.3%
CaseLaw v263.1%
Code Migration43.0%
Excel Modeling Benchmark62.9%
Finance Agent v258.3%
Harvey's Legal Agent Benchmark10.8%
Legal Research Bench47.1%
MedCode51.3%
MedScribe85.9%
MMMU-Pro (Vals)87.4%
Terminal-Bench 2.0 (Vals)59.5%
Terminal-Bench Science2.9%
Vibe Code Bench v1.185.2%19.7%
FORTRESS20.2%
SWE-Bench Pro (private)55.0%
SWE Atlas: Codebase QnA24.2%
SWE Atlas: Test Writing31.1%
LMArena Maths1461
LMArena Creative Writing1464
LMArena Instruction Following1463
LMArena Multi-turn1491
LMArena Longer Queries1474
LMArena Document1467
Terminal-Bench 4.0 (AA)34.9%
Terminal-Bench 2.1 (AA)62.2%
AutomationBench58.6%
GDP.pdf19.2%
AA-Omniscience: accuracy34.9%49.6%
AA-Omniscience: non-hallucination59.4%15.8%
AA-Briefcase1522

Data as of 2026-09-23. Best configuration of each model; every score links to its source on the model pages.

MiMo-V2.6-Pro vs Muse Spark: questions

Is MiMo-V2.6-Pro better than Muse Spark?
Muse Spark leads on quality: 65.3 vs 64.3. The BenchLeader Index combines every independent quality benchmark; Muse Spark is ahead overall as of 2026-09-23, but check the category scores for your use.
Is MiMo-V2.6-Pro better than Muse Spark for agentic tasks?
Muse Spark scores higher in agentic tasks (66 vs 55 on the category index, where 50 is average).
Which has the larger context window?
MiMo-V2.6-Pro accepts more context: 1.0M against 262k tokens.