BenchLeader
MiniMaxOpen weights: MiniMaxAI/Minimax-M3Reasoning modelAuto-detected

MiniMax-M3

Best configuration ranks #152 of 610 on the BenchLeader Index at 56.9 ±5.0. Last measured 8 Sept 2026. Released 1 Jun 2026.

Blended price
$0.525/M
$0.300 in · $1.20 out
Output speed
100 tok/s
First answer
22 s
first token 1.12 s
Context
1.0M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index57
  2. Reasoning58
  3. Coding53
  4. Agents & tools45
  5. Maths48
  6. Knowledge60
  7. Instruction following79
  8. Human preference57
  9. Multimodal62
  10. Long context68
  11. Composite47

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning100 tok/s22 s$0.0005Maths 36 · Reasoning 59
defaultbest56.9#152100 tok/s22 s$0.0005Agents & tools 45 · Coding 53 · Composite 47 · Human preference 57 · Instruction following 79 · Knowledge 60 · Long context 68 · Maths 48 · Multimodal 62 · Reasoning 58

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoningdefaultSourceTrend
GPQA Diamond81.3%#10890.9%#29Epoch AI Benchmarking Hub
SimpleBench45.8%#52SimpleBench
LMArena Hard Prompts1464#73LMArena
LiveBench Reasoningnot in index74.5%#48LiveBench
GPQA Diamond (AA)not in index92.9%#27Artificial Analysis
Humanity's Last Exam (AA)not in index39.0%#61Artificial Analysis
GPQA Diamond (Vals)not in index92.7%#14Vals AI

Coding

Benchmarkno reasoningdefaultSourceTrend
SciCode45.4%#85SciCode
FrontierCode14.7%#26Cognition
LMArena Coding1498#66LMArena
LMArena WebDev1486#44LMArena
LiveBench Codingnot in index68.2%#50LiveBench
SciCode (AA)not in index47.1%#83Artificial Analysis
LiveCodeBench82.2%#54Vals AI
SWE-bench (Vals)not in index75.0%#42Vals AI

Agents & tools

Maths

Benchmarkno reasoningdefaultSourceTrend
OTIS Mock AIME26.7%#19371.1%#120Epoch AI Benchmarking Hub
ProofBench18.0%#38Vals AI
LiveBench Mathematicsnot in index77.0%#51LiveBench

Knowledge

Benchmarkno reasoningdefaultSourceTrend
LiveBench Data Analysisnot in index76.2%#25LiveBench
AA-Omniscience1.4#84Artificial Analysis
MMLU-Pro84.2%#63Vals AI
LegalBench85.4%#18Vals AI
CorpFin68.1%#10Vals AI
TaxEval72.7%#58Vals AI

Instruction following

Benchmarkno reasoningdefaultSourceTrend
LiveBench Languagenot in index76.8%#34LiveBench
IFBench82.9%#3Artificial Analysis

Human preference

Benchmarkno reasoningdefaultSourceTrend
LMArena Text1443#77LMArena
EQ-Bench 41150#19EQ-Bench

Multimodal

Benchmarkno reasoningdefaultSourceTrend
LMArena Vision1255#44LMArena
MMMU-Pro78.5%#48Artificial Analysis

Long context

Benchmarkno reasoningdefaultSourceTrend
AA-LCR83.0%#12Artificial Analysis

Composite

Benchmarkno reasoningdefaultSourceTrend
Epoch Capabilities Indexnot in index146.5#61Epoch AI Benchmarking Hub
LiveBench67.3%#48LiveBench
AA Intelligence Index29.6#83Artificial Analysis
Vals Indexnot in index42.7#33Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
CoreWeave142 tok/s0.44 s$0.230$0.960262kfp4
SambaNova134 tok/s1.47 s$0.600$2.401.0M
ModelRun [by Modular]102 tok/s1.15 s$0.750$3.001.0Mfp4
Venice83 tok/s1.05 s$0.300$1.20524kfp8
MiniMax77 tok/s0.96 s$0.300$1.20524kfp8
Together66 tok/s0.90 s$0.300$1.20524k
Parasail66 tok/s0.87 s$0.300$1.201.0Mfp8
NovitaAI63 tok/s1.63 s$0.300$1.201Mfp8
GMICloud59 tok/s1.33 s$0.240$0.9601.0Mfp8
DeepInfra27 tok/s1.12 s$0.280$1.10524kfp8
StreamLake26 tok/s1.44 s$0.300$1.201Mfp8

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.660$1.32$1.98$2.64Jul 26Aug 26
input outputnow $0.240 in · $0.960 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.060 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0005$0.000424.7 s
Summarise a 30-page report12,000 / 600$0.0043$0.002227.7 s
Code edit6,000 / 1,500$0.0036$0.002536.6 s
Agentic coding session60,000 / 4,000$0.023$0.0121.0 min
Structured extraction2,000 / 200$0.0008$0.000523.7 s

See also

Data as of 9 Sept 2026. Compare these configurations.