BenchLeader

MiMo-V2.6-Pro vs Qwen3.8 Max (0902)

Verdict
  • Qwen3.8 Max (0902) (max) leads on quality: 65.9 vs 64.3.
  • MiMo-V2.6-Pro is stronger in composite, knowledge, long context, reasoning.
  • Qwen3.8 Max (0902) (max) is stronger in agents & tools, coding, human preference, maths, multimodal.
  • MiMo-V2.6-Pro is 5.5× cheaper ($0.548 vs $3.00 per 1M blended).
  • MiMo-V2.6-Pro streams 1.4× faster (54 vs 39 tokens per second).
MetricMiMo-V2.6-ProQwen3.8 Max (0902) (max)
BenchLeader Index64.365.9
Agents & tools score55.168.0
Composite score84.872.1
Knowledge score66.863.1
Long context score69.166.0
Reasoning score88.874.1
Coding score66.1
Human preference score68.3
Maths score67.1
Multimodal score67.0
Blended price $/M$0.548$3.00
Output speed54 tok/s39 tok/s
Time to first answer39.8 s55.1 s
Context window1.0M1M
SciCode52.9%
ProofBench58.0%
Epoch Capabilities Index156.7
LMArena Text1481
LMArena Hard Prompts1503
LMArena Coding1522
LMArena WebDev1671
LMArena Vision1315
LMArena Agent3.3
LiveBench78.5%
LiveBench Reasoning88.2%
LiveBench Coding72.9%
LiveBench Agentic Coding64.7%
LiveBench Mathematics91.3%
LiveBench Data Analysis78.4%
LiveBench Language79.7%
LiveBench Instruction Following74.1%
AA Intelligence Index46.345.4
AA-LCR86.3%80.3%
MMMU-Pro82.8%
AA-Omniscience8.412.0
GPQA Diamond (AA)92.8%
Humanity's Last Exam (AA)49.4%43.1%
SciCode (AA)60.9%53.2%
LiveCodeBench87.8%
MMLU-Pro88.6%
IOI68.9%
LegalBench83.6%
CorpFin65.8%
TaxEval75.5%
Terminal-Bench 2.1 (Vals)67.8%67.4%
SWE-bench (Vals)85.6%
GPQA Diamond (Vals)93.7%
Vals Index59.751.8
CritPt26.6%20.0%
GDPval (AA)58.7%58.4%
τ²-Bench Banking (AA)51.3%
Code Migration43.0%24.0%
Excel Modeling Benchmark62.9%60.1%
Finance Agent v258.3%50.6%
Harvey's Legal Agent Benchmark10.8%10.4%
Legal Research Bench47.1%47.6%
MedCode40.7%
MedScribe85.0%
MMMU-Pro (Vals)88.0%
MortgageTax64.0%
MysteryMechanism23.9%
ProgramBench0.0%
Public Benefits Bench67.1%
SAGE51.3%
SkillsBench42.0%
Tax Agent Bench66.0%
Terminal-Bench 4.0 (Vals)24.8%
Terminal-Bench Science2.9%1.4%
Vals Multimodal Index65.4%
Vibe Code Bench 1-10012.8%
Vibe Code Bench v1.185.2%64.7%
LMArena Maths1498
LMArena Creative Writing1472
LMArena Instruction Following1477
LMArena Multi-turn1495
LMArena Longer Queries1492
Terminal-Bench 4.0 (AA)34.9%38.9%
Terminal-Bench 2.1 (AA)88.8%
AutomationBench58.6%56.2%
GDP.pdf19.2%22.8%
MLCR20.0%
AA-Omniscience: accuracy34.9%31.9%
AA-Omniscience: non-hallucination59.4%71.2%
AA-Briefcase15221640

Data as of 2026-09-23. Best configuration of each model; every score links to its source on the model pages.

MiMo-V2.6-Pro vs Qwen3.8 Max (0902): questions

Is MiMo-V2.6-Pro better than Qwen3.8 Max (0902)?
Qwen3.8 Max (0902) (max) leads on quality: 65.9 vs 64.3. The BenchLeader Index combines every independent quality benchmark; Qwen3.8 Max (0902) (max) is ahead overall as of 2026-09-23, but check the category scores for your use.
Is MiMo-V2.6-Pro better than Qwen3.8 Max (0902) for agentic tasks?
Qwen3.8 Max (0902) scores higher in agentic tasks (68 vs 55 on the category index, where 50 is average).
Which is cheaper, MiMo-V2.6-Pro or Qwen3.8 Max (0902)?
MiMo-V2.6-Pro is cheaper: $0.548 against $3.00 per million tokens, blended at three input tokens per output token.
Which is faster, MiMo-V2.6-Pro or Qwen3.8 Max (0902)?
MiMo-V2.6-Pro streams faster: 54 against 39 output tokens per second.
Which has the larger context window?
MiMo-V2.6-Pro accepts more context: 1.0M against 1M tokens.