BenchLeader

Grok 4.6 vs Mimo v2 Pro

Verdict
  • Grok 4.6 (medium) leads on quality: 63.9 vs 60.6.
  • Grok 4.6 (medium) is stronger in coding, composite, knowledge, long context.
  • Mimo v2 Pro is stronger in reasoning, agents & tools, human preference, instruction following.
  • They cost about the same ($3.00 per 1M blended).
MetricGrok 4.6 (medium)Mimo v2 Pro
BenchLeader Index63.960.6
Coding score64.154.7
Composite score83.465.2
Knowledge score77.566.3
Long context score67.160.5
Reasoning score62.165.2
Agents & tools score69.4
Human preference score64.5
Instruction following score67.5
Blended price $/M$3.00$3.00
Output speed53 tok/s
Time to first answer33.4 s
Context window500k1.0M
SciCode54.6%
LMArena Text1448
LMArena Hard Prompts1476
LMArena Coding1503
LMArena WebDev1433
AA Intelligence Index43.028.6
IFBench68.8%
AA-LCR81.0%68.3%
AA-Omniscience284.6
Terminal-Bench Hard40.9%
GPQA Diamond (AA)93.5%87.0%
Humanity's Last Exam (AA)42.1%30.4%
SciCode (AA)55.9%
τ²-Bench Telecom (AA)95.0%
ARC-AGI-187.5%
ARC-AGI-261.3%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Grok 4.6 vs Mimo v2 Pro: questions

Is Grok 4.6 better than Mimo v2 Pro?
Grok 4.6 (medium) leads on quality: 63.9 vs 60.6. The BenchLeader Index combines every independent quality benchmark; Grok 4.6 (medium) is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Grok 4.6 better than Mimo v2 Pro for coding?
Grok 4.6 scores higher in coding (64 vs 55 on the category index, where 50 is average).
Which is cheaper, Grok 4.6 or Mimo v2 Pro?
Grok 4.6 is cheaper: $3.00 against $3.00 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Mimo v2 Pro accepts more context: 1.0M against 500k tokens.