Mistral Small 2402
Mistral Small 2402 is a Mistral AI proprietary model. Its best configuration ranks #623 of 372 on the BenchLeader Index at 32.0 ±7.5, in the lower half. It scores highest in composite (36) and lowest in coding (7). At $0.262 per million tokens blended it is cheaper than most ranked models. Output speed of 143 tokens per second puts it in the fastest quarter, with a first answer in 0.8 s. Last measured 1 Sept 2026.
- Blended price
- $0.262/M
- $0.150 in · $0.600 out
- Output speed
- 143 tok/s
- measured by Artificial Analysis
- First answer
- 0.76 s
- first token 0.76 s
- Context
- 33k
- Overall index32
- Coding7
- Maths32
- Knowledge22
- Composite36
Versions
Mistral AI has shipped 8 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.
| Model | Released | Index | Rank |
|---|---|---|---|
| Mistral Small 4 | 16 Mar 2026 | 47.2 | #364 |
| Mistral Small 3.1 | 17 Mar 2025 | 39.8 | #535 |
| Mistral Small 3.1 | 17 Mar 2025 | 34.6 | #616 |
| Mistral Small 3 | 30 Jan 2025 | 38.4 | #565 |
| Mistral Small 3 | 25 Jan 2025 | – | – |
| Mistral Small 3.2 | – | 42.3 | #474 |
| Mistral Small 2402this page | – | 32.0 | #623 |
| Mistral Small 3 | – | – | – |
Benchmark results
One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.
Reasoning
| Benchmark | default | Source |
|---|---|---|
| GPQA Diamond (AA)not in index | 30.2%#501 | Artificial Analysis |
| Humanity's Last Exam (AA)not in index | 4.1%#468 | Artificial Analysis |
| GPQA Diamond (Vals)not in index | 50.8%#114 | Vals AI |
Coding
| Benchmark | default | Source |
|---|---|---|
| LiveCodeBench | 15.8%#137 | Vals AI |
Maths
| Benchmark | default | Source |
|---|---|---|
| AIME (Vals) | 5.6%#88 | Vals AI |
| MGSM | 84.0%#67 | Vals AI |
Composite
| Benchmark | default | Source |
|---|---|---|
| AA Intelligence Index | 5.5#502 | Artificial Analysis |
What a task costs
Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.
| Workload | Tokens in / out | Cost | With caching | Time |
|---|---|---|---|---|
| Chat reply | 400 / 300 | $0.0002 | – | 2.9 s |
| Summarise a 30-page report | 12,000 / 600 | $0.0022 | – | 4.9 s |
| Code edit | 6,000 / 1,500 | $0.0018 | – | 11.2 s |
| Agentic coding session | 60,000 / 4,000 | $0.011 | – | 28.7 s |
| Structured extraction | 2,000 / 200 | $0.0004 | – | 2.2 s |
See also
Data as of 10 Sept 2026. Compare with another model.