BenchLeader
Mistral AIAuto-detected

Mistral Small 2402

Mistral Small 2402 is a Mistral AI proprietary model. Its best configuration ranks #623 of 372 on the BenchLeader Index at 32.0 ±7.5, in the lower half. It scores highest in composite (36) and lowest in coding (7). At $0.262 per million tokens blended it is cheaper than most ranked models. Output speed of 143 tokens per second puts it in the fastest quarter, with a first answer in 0.8 s. Last measured 1 Sept 2026.

Blended price
$0.262/M
$0.150 in · $0.600 out
Output speed
143 tok/s
measured by Artificial Analysis
First answer
0.76 s
first token 0.76 s
Context
33k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index32
  2. Coding7
  3. Maths32
  4. Knowledge22
  5. Composite36

Versions

Mistral AI has shipped 8 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
Mistral Small 416 Mar 202647.2#364
Mistral Small 3.117 Mar 202539.8#535
Mistral Small 3.117 Mar 202534.6#616
Mistral Small 330 Jan 202538.4#565
Mistral Small 325 Jan 2025
Mistral Small 3.242.3#474
Mistral Small 2402this page32.0#623
Mistral Small 3

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkdefaultSource
GPQA Diamond (AA)not in index30.2%#501Artificial Analysis
Humanity's Last Exam (AA)not in index4.1%#468Artificial Analysis
GPQA Diamond (Vals)not in index50.8%#114Vals AI

Coding

BenchmarkdefaultSource
LiveCodeBench15.8%#137Vals AI

Maths

BenchmarkdefaultSource
AIME (Vals)5.6%#88Vals AI
MGSM84.0%#67Vals AI

Knowledge

BenchmarkdefaultSource
MMLU-Pro64.4%#124Vals AI
TaxEval49.1%#134Vals AI
MedQA57.0%#84Vals AI

Composite

BenchmarkdefaultSource
AA Intelligence Index5.5#502Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.00022.9 s
Summarise a 30-page report12,000 / 600$0.00224.9 s
Code edit6,000 / 1,500$0.001811.2 s
Agentic coding session60,000 / 4,000$0.01128.7 s
Structured extraction2,000 / 200$0.00042.2 s

See also

Data as of 10 Sept 2026. Compare with another model.