BenchLeader

MiniMax-M2.5

Best configuration ranks #220 of 610 on the BenchLeader Index at 53.5 ±4.4. Last measured 8 Sept 2026. Released 12 Feb 2026.

Blended price
$0.525/M
$0.300 in · $1.20 out
Output speed
93 tok/s
First answer
23 s
first token 1.04 s
Context
205k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index54
  2. Reasoning47
  3. Coding50
  4. Agents & tools44
  5. Maths48
  6. Knowledge50
  7. Instruction following70
  8. Human preference57
  9. Long context63
  10. Composite58

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
high93 tok/s23 s$0.0005Coding 66
defaultbest53.5#22093 tok/s23 s$0.0005Agents & tools 44 · Coding 50 · Composite 58 · Human preference 57 · Instruction following 70 · Knowledge 50 · Long context 63 · Maths 48 · Reasoning 47

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkhighdefaultSource
LMArena Hard Prompts1416#138LMArena
GPQA Diamond (AA)not in index84.8%#131Artificial Analysis
Humanity's Last Exam (AA)not in index20.5%#158Artificial Analysis
GPQA Diamond (Vals)not in index82.1%#60Vals AI
Kagi LLM Benchmark55.2%#64Kagi LLM Benchmark
ARC-AGI-163.7%#104ARC Prize
ARC-AGI-24.9%#120ARC Prize

Coding

BenchmarkhighdefaultSource
LMArena Coding1444#133LMArena
LMArena WebDev1384#79LMArena
LiveCodeBench79.2%#70Vals AI
IOI6.7%#39Vals AI
SWE-bench (Vals)not in index74.2%#47Vals AI
SWE-bench Verified (bash only)75.8%#2SWE-bench
SWE-bench Verified (any scaffold)not in index75.8%#5SWE-bench

Agents & tools

Maths

BenchmarkhighdefaultSource
ProofBench4.0%#56Vals AI
AIME (Vals)88.8%#28Vals AI

Knowledge

BenchmarkhighdefaultSource
AA-Omniscience-38.9#251Artificial Analysis
MMLU-Pro80.1%#86Vals AI
LegalBench80.0%#84Vals AI
CorpFin59.6%#74Vals AI
TaxEval68.2%#97Vals AI
MedQA92.5%#29Vals AI

Instruction following

BenchmarkhighdefaultSource
IFBench71.6%#52Artificial Analysis

Human preference

BenchmarkhighdefaultSource
LMArena Text1391#141LMArena

Long context

BenchmarkhighdefaultSource
AA-LCR73.3%#129Artificial Analysis

Composite

BenchmarkhighdefaultSource
Epoch Capabilities Indexnot in index146.5#59Epoch AI Benchmarking Hub
AA Intelligence Index22.8#140Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Friendli128 tok/s0.44 s$0.300$1.20197k
Venice68 tok/s1.04 s$0.270$0.950198k
StreamLake66 tok/s1.04 s$0.270$1.08200k
DigitalOcean64 tok/s0.51 s$0.300$1.2066k
NovitaAI55 tok/s1.34 s$0.300$1.20205kfp8
MiniMax Highspeed52 tok/s1.00 s$0.600$2.40205kfp8
MiniMax50 tok/s1.12 s$0.300$1.20205kfp8
AtlasCloud40 tok/s2.51 s$0.295$1.20197kfp8

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.330$0.660$0.990$1.32Jun 26Jul 26Aug 26
input outputnow $0.300 in · $1.20 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.030 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0005$0.000426.5 s
Summarise a 30-page report12,000 / 600$0.0043$0.001929.7 s
Code edit6,000 / 1,500$0.0036$0.002439.4 s
Agentic coding session60,000 / 4,000$0.023$0.0111.1 min
Structured extraction2,000 / 200$0.0008$0.000425.4 s

See also

Data as of 9 Sept 2026. Compare these configurations.