BenchLeader

MiMo-V2.5-Pro vs Qwen3 8

Verdict
  • Qwen3 8 (max) leads on quality: 66.5 vs 59.2.
  • MiMo-V2.5-Pro is stronger in instruction following.
  • Qwen3 8 (max) is stronger in agents & tools, coding, composite, human preference, knowledge, long context, maths, reasoning, multimodal.
  • MiMo-V2.5-Pro is 4.1× cheaper ($0.653 vs $2.67 per 1M blended).
  • MiMo-V2.5-Pro streams 1.4× faster (53 vs 39 tokens per second).
MetricMiMo-V2.5-ProQwen3 8 (max)
BenchLeader Index59.266.5
Agents & tools score45.968.3
Coding score58.366.5
Composite score62.574.3
Human preference score61.568.2
Instruction following score77.3
Knowledge score57.663.5
Long context score66.366.6
Maths score59.067.0
Reasoning score55.676.3
Multimodal score67.4
Blended price $/M$0.653$2.67
Output speed53 tok/s39 tok/s
Time to first answer42.2 s55.1 s
Context window1.0M1M
SciCode50.2%52.9%
ProofBench22.0%58.0%
Epoch Capabilities Index156.6
LMArena Text14671481
LMArena Hard Prompts14961503
LMArena Coding15211522
LMArena WebDev14751671
LMArena Vision1315
LMArena Agent-5.73.3
LiveBench78.5%
LiveBench Reasoning88.2%
LiveBench Coding72.9%
LiveBench Agentic Coding64.7%
LiveBench Mathematics91.3%
LiveBench Data Analysis78.4%
LiveBench Language79.7%
LiveBench Instruction Following74.1%
AA Intelligence Index26.445.4
IFBench79.9%
AA-LCR79.7%80.3%
MMMU-Pro82.8%
AA-Omniscience3.312.0
Terminal-Bench Hard43.2%
GPQA Diamond (AA)86.6%92.8%
Humanity's Last Exam (AA)35.7%43.1%
SciCode (AA)50.6%53.2%
τ²-Bench Telecom (AA)94.2%
LiveCodeBench81.3%87.8%
MMLU-Pro84.6%88.6%
IOI68.9%
LegalBench77.1%83.6%
CorpFin61.4%65.8%
TaxEval73.8%75.5%
Terminal-Bench 2.1 (Vals)57.3%67.4%
SWE-bench (Vals)74.0%85.6%
GPQA Diamond (Vals)82.6%93.7%
Vals Index41.051.8
EQ-Bench 41208
CritPt4.0%20.0%
GDPval (AA)34.3%58.2%
τ²-Bench Banking (AA)9.9%51.3%
ITBench SRE (AA)38.2%
Analyst Agent (AA)20.0%
APEX-Agents (AA)2.4%
Code Migration21.6%24.0%
Excel Modeling Benchmark55.2%60.1%
Finance Agent v241.5%50.6%
Harvey's Legal Agent Benchmark2.1%10.4%
Legal Research Bench15.9%47.6%
MedCode32.5%40.7%
MedScribe83.7%85.0%
MMMU-Pro (Vals)88.0%
MortgageTax64.0%
MysteryMechanism23.9%
ProgramBench0.0%
Public Benefits Bench67.1%
SAGE51.3%
SkillsBench42.0%
Tax Agent Bench66.0%
Terminal-Bench 4.0 (Vals)24.8%
Terminal-Bench Science1.4%
Vals Multimodal Index65.4%
Vibe Code Bench 1-10012.8%
Vibe Code Bench v1.134.1%64.7%
LMArena Maths14761498
LMArena Creative Writing14351472
LMArena Instruction Following14701477
LMArena Multi-turn14781495
LMArena Longer Queries14871492
LMCA29.5%
DTBench84.5%
ALE-Bench899.8

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

MiMo-V2.5-Pro vs Qwen3 8: questions

Is MiMo-V2.5-Pro better than Qwen3 8?
Qwen3 8 (max) leads on quality: 66.5 vs 59.2. The BenchLeader Index combines every independent quality benchmark; Qwen3 8 (max) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is MiMo-V2.5-Pro better than Qwen3 8 for coding?
Qwen3 8 scores higher in coding (67 vs 58 on the category index, where 50 is average).
Is MiMo-V2.5-Pro better than Qwen3 8 for agentic tasks?
Qwen3 8 scores higher in agentic tasks (68 vs 46 on the category index, where 50 is average).
Which is cheaper, MiMo-V2.5-Pro or Qwen3 8?
MiMo-V2.5-Pro is cheaper: $0.653 against $2.67 per million tokens, blended at three input tokens per output token.
Which is faster, MiMo-V2.5-Pro or Qwen3 8?
MiMo-V2.5-Pro streams faster: 53 against 39 output tokens per second.
Which has the larger context window?
MiMo-V2.5-Pro accepts more context: 1.0M against 1M tokens.