BenchLeader

DeepSeek V4 Pro vs MiMo-V2.6-Pro

Verdict
  • DeepSeek V4 Pro (max) and MiMo-V2.6-Pro are level on quality (63.4 vs 64.3).
  • DeepSeek V4 Pro (max) is stronger in agents & tools, coding, instruction following, maths.
  • MiMo-V2.6-Pro is stronger in composite, knowledge, long context, reasoning.
  • They cost about the same ($0.544 per 1M blended).
  • DeepSeek V4 Pro (max) streams 1.3× faster (71 vs 54 tokens per second).
MetricDeepSeek V4 Pro (max)MiMo-V2.6-Pro
BenchLeader Index63.464.3
Agents & tools score63.455.1
Coding score58.9
Composite score72.484.8
Instruction following score74.3
Knowledge score60.266.8
Long context score66.069.1
Maths score57.1
Reasoning score66.688.8
Blended price $/M$0.544$0.548
Output speed71 tok/s54 tok/s
Time to first answer63.0 s39.8 s
Context window1M1.0M
GPQA Diamond91.7%
FrontierMath Tiers 1–364.6%
FrontierMath Tier 426.8%
OTIS Mock AIME98.6%
SWE-bench Verified (Epoch)77.6%
SimpleQA Verified52.9%
SciCode51.0%
WeirdML66.2%
ProofBench16.0%
AA Intelligence Index3646.3
IFBench76.5%
AA-LCR80.3%86.3%
AA-Omniscience0.88.4
Terminal-Bench Hard46.2%
GPQA Diamond (AA)92.8%
Humanity's Last Exam (AA)41.0%49.4%
SciCode (AA)51.0%60.9%
τ²-Bench Telecom (AA)96.2%
LiveCodeBench87.5%
MMLU-Pro87.3%
IOI51.6%
LegalBench82.4%
CorpFin65.4%
TaxEval73.1%
Terminal-Bench 2.1 (Vals)54.7%67.8%
SWE-bench (Vals)96.4%
GPQA Diamond (Vals)92.4%
Vals Index52.459.7
AIME 202696.7%
HMMT February 202693.9%
MathArena Apex28.1%
ARC-AGI-190.0%
ARC-AGI-261.3%
CritPt18.0%26.6%
GDPval (AA)47.1%58.7%
τ²-Bench Banking (AA)39.6%
ITBench SRE (AA)38.3%
Analyst Agent (AA)18.8%
APEX-Agents (AA)24.3%
CaseLaw v259.4%
Code Migration41.5%43.0%
Excel Modeling Benchmark52.8%62.9%
Finance Agent v250.4%58.3%
Harvey's Legal Agent Benchmark7.5%10.8%
Legal Research Bench40.9%47.1%
MedCode42.5%
MedScribe80.2%
ProgramBench0.0%
Public Benefits Bench62.9%
SkillsBench53.8%
Tax Agent Bench58.7%
Terminal-Bench 4.0 (Vals)1.0%
Terminal-Bench Science0.0%2.9%
Vibe Code Bench 1-10017.5%
Vibe Code Bench v1.182.3%85.2%
Chess Puzzles47.0%
Mystery Game Puzzles43.0%
LMCA41.2%
DTBench90.7%
ALE-Bench1403.2
Terminal-Bench 4.0 (AA)14.7%34.9%
Terminal-Bench 2.1 (AA)78.7%
AutomationBench56.7%58.6%
GDP.pdf11.4%19.2%
MLCR17.8%
EnterpriseOps-Gym49.6%
AA-Omniscience: accuracy49.1%34.9%
AA-Omniscience: non-hallucination5.9%59.4%
AA-Briefcase12611522
AA Openness Index44.4

Data as of 2026-09-23. Best configuration of each model; every score links to its source on the model pages.

DeepSeek V4 Pro vs MiMo-V2.6-Pro: questions

Is DeepSeek V4 Pro better than MiMo-V2.6-Pro?
DeepSeek V4 Pro (max) and MiMo-V2.6-Pro are level on quality (63.4 vs 64.3). The BenchLeader Index combines every independent quality benchmark; MiMo-V2.6-Pro is ahead overall as of 2026-09-23, but check the category scores for your use.
Is DeepSeek V4 Pro better than MiMo-V2.6-Pro for agentic tasks?
DeepSeek V4 Pro scores higher in agentic tasks (63 vs 55 on the category index, where 50 is average).
Which is cheaper, DeepSeek V4 Pro or MiMo-V2.6-Pro?
DeepSeek V4 Pro is cheaper: $0.544 against $0.548 per million tokens, blended at three input tokens per output token.
Which is faster, DeepSeek V4 Pro or MiMo-V2.6-Pro?
DeepSeek V4 Pro streams faster: 71 against 54 output tokens per second.
Which has the larger context window?
MiMo-V2.6-Pro accepts more context: 1.0M against 1M tokens.