BenchLeader

MiMo-V2.6-Flash vs Muse Spark 1.3

Verdict
  • Muse Spark 1.3 (max) leads on quality: 67.7 vs 59.7.
  • Muse Spark 1.3 (max) is stronger in agents & tools, coding, composite, human preference, knowledge, long context, maths, multimodal, reasoning.
  • MiMo-V2.6-Flash is 11× cheaper ($0.175 vs $2.00 per 1M blended).
  • Muse Spark 1.3 (max) streams 3.0× faster (175 vs 58 tokens per second).
MetricMiMo-V2.6-FlashMuse Spark 1.3 (max)
BenchLeader Index59.767.7
Agents & tools score55.570.3
Coding score60.164.6
Composite score72.584.1
Human preference score64.069.4
Knowledge score56.072.9
Long context score62.366.8
Maths score63.463.6
Multimodal score57.865.9
Reasoning score62.578.1
Blended price $/M$0.175$2.00
Output speed58 tok/s175 tok/s
Time to first answer38.1 s70.5 s
Context window1.0M1.0M
FrontierMath Tiers 1–3–74.0%
FrontierMath Tier 4–46.3%
SciCode51.3%58.8%
ProofBench63.0%58.0%
LMArena Text14511494
LMArena Hard Prompts14841517
LMArena Coding15141538
LMArena WebDev16371657
LMArena Vision12591309
LMArena Agent0.44
AA Intelligence Index v4.3.237.948.1
AA-LCR74.3%83.0%
MMMU-Pro73.1%–
AA-Omniscience-12.725
GPQA Diamond (AA)–93.5%
Humanity's Last Exam (AA)35.1%48.7%
SciCode (AA)51.3%58.8%
IOI47.7%56.6%
Terminal-Bench 2.1 (Vals)76.4%79.0%
Vals Index53.258.2
CritPt12.0%24.9%
GDPval-AA v2.155.5%59.0%
τ³-Banking (AA)–50.5%
ITBench SRE (AA)–33.2%
BioMysteryBench69.3%–
Code Migration40.9%47.4%
CUA-bench–5.8%
CyberBench75.4%72.7%
Excel Modeling Benchmark65.5%67.4%
Finance Agent v256.3%60.0%
Harvey's Legal Agent Benchmark11.3%23.8%
Legal Research Bench38.0%55.3%
MedCode41.1%–
MedScribe85.3%–
MysteryMechanism21.6%36.0%
ProgramBench0.5%–
Public Benefits Bench67.6%–
SAGE43.5%–
SREBench4.2%–
Tax Agent Bench59.9%72.4%
Terminal-Bench 4.0 (Vals)24.2%24.8%
Terminal-Bench Science4.3%10.0%
Vibe Code Bench 1-100–20.5%
Vibe Code Bench v1.179.0%85.9%
LMArena Maths14621502
LMArena Creative Writing13871456
LMArena Instruction Following14511483
LMArena Multi-turn14521487
LMArena Longer Queries14631497
LMArena Document–1468
Chess Puzzles–38.0%
Mystery Game Puzzles–25.0%
CursorBench–41.6%
Terminal-Bench 4.0 (AA)22.7%33.3%
Terminal-Bench 2.1 (AA)–84.3%
AutomationBench–57.9%
GDP.pdf–26.6%
MLCR–43.3%
Harvey LAB–8.9%
AA-Omniscience: accuracy27.0%43.6%
AA-Omniscience: non-hallucination45.6%67.1%
AA-Briefcase v1.1–1581

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

MiMo-V2.6-Flash vs Muse Spark 1.3: questions

Is MiMo-V2.6-Flash better than Muse Spark 1.3?
Muse Spark 1.3 (max) leads on quality: 67.7 vs 59.7. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is MiMo-V2.6-Flash better than Muse Spark 1.3 for coding?
Muse Spark 1.3 scores higher in coding (65 vs 60 on the category index, where 50 is average).
Is MiMo-V2.6-Flash better than Muse Spark 1.3 for agentic tasks?
Muse Spark 1.3 scores higher in agentic tasks (70 vs 56 on the category index, where 50 is average).
Which is cheaper, MiMo-V2.6-Flash or Muse Spark 1.3?
MiMo-V2.6-Flash is cheaper: $0.175 against $2.00 per million tokens, blended at three input tokens per output token.
Which is faster, MiMo-V2.6-Flash or Muse Spark 1.3?
Muse Spark 1.3 streams faster: 175 against 58 output tokens per second.
Which has the larger context window?
Both accept 1.0M tokens of context.