BenchLeader

MiMo-V2-Flash

Best configuration ranks #224 of 610 on the BenchLeader Index at 53.4 ±5.8 (thinking reasoning effort). Last measured 8 Sept 2026. Released 16 Dec 2025.

Blended price
$0.150/M
$0.100 in · $0.300 out
Output speed
First answer
Context
256k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index53
  2. Reasoning58
  3. Coding42
  4. Agents & tools58
  5. Knowledge43
  6. Instruction following63
  7. Human preference57
  8. Long context62
  9. Composite55

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning$0.0001Coding 45 · Human preference 57 · Reasoning 58
thinkingbest53.4#224$0.0001Agents & tools 58 · Coding 42 · Composite 55 · Human preference 57 · Instruction following 63 · Knowledge 43 · Long context 62 · Reasoning 58
default53.3#228$0.0001Agents & tools 61 · Coding 25 · Composite 57 · Instruction following 70 · Knowledge 55 · Long context 62

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoningthinkingdefaultSourceTrend
LMArena Hard Prompts1414#1391411#144LMArena
GPQA Diamond (AA)not in index84.7%#13483.5%#154Artificial Analysis
Humanity's Last Exam (AA)not in index22.9%#14522.1%#149Artificial Analysis
GPQA Diamond (Vals)not in index59.3%#106Vals AI

Coding

Benchmarkno reasoningthinkingdefaultSourceTrend
SciCode25.9%#148SciCode
LMArena Coding1446#1301430#153LMArena
LMArena WebDev1331#961293#100LMArena

Agents & tools

Benchmarkno reasoningthinkingdefaultSourceTrend
Terminal-Bench Hard28.0%#12231.1%#104Artificial Analysis
τ²-Bench Telecom (AA)not in index95.0%#2793.3%#44Artificial Analysis

Knowledge

Benchmarkno reasoningthinkingdefaultSourceTrend
AA-Omniscience-45.4#285-19.1#179Artificial Analysis

Instruction following

Benchmarkno reasoningthinkingdefaultSourceTrend
IFBench64.2%#11071.8%#51Artificial Analysis

Human preference

Benchmarkno reasoningthinkingdefaultSourceTrend
LMArena Text1392#1391386#147LMArena

Long context

Benchmarkno reasoningthinkingdefaultSourceTrend
AA-LCR70.7%#15471.3%#149Artificial Analysis

Composite

Benchmarkno reasoningthinkingdefaultSourceTrend
AA Intelligence Index20.8#16622.4#144Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.300 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0001$0.0002
Summarise a 30-page report12,000 / 600$0.0014$0.0032
Code edit6,000 / 1,500$0.0011$0.0019
Agentic coding session60,000 / 4,000$0.0072$0.016
Structured extraction2,000 / 200$0.0003$0.0006

See also

Data as of 9 Sept 2026. Compare these configurations.