BenchLeader
Mistral AIAuto-detected

Magistral Small 1.2

Magistral Small 1.2 is a Mistral AI proprietary model, released 18 Sept 2025. Its best configuration ranks #505 of 372 on the BenchLeader Index at 41.2 ±5.6, in the lower half. It scores highest in coding (47) and lowest in knowledge (28). At $0.750 per million tokens blended it is mid-priced. Last measured 10 Sept 2026.

Blended price
$0.750/M
$0.500 in · $1.50 out
Output speed
First answer
Context
128k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index41
  2. Reasoning37
  3. Coding47
  4. Agents & tools38
  5. Maths47
  6. Knowledge28
  7. Instruction following46
  8. Multimodal39
  9. Long context35
  10. Composite40

Versions

Mistral AI has shipped 5 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
Magistral Small 1.2this page18 Sept 202541.2#505
Magistral Small 1.010 Jun 2025
Magistral Small17 Mar 202537.9#570
Magistral Small 1.0
Magistral Small 1.2

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Coding

BenchmarkdefaultSource
SciCode35.2%#137SciCode
LiveCodeBench72.1%#82Vals AI

Agents & tools

Maths

BenchmarkdefaultSource
OTIS Mock AIME28.1%#197Epoch AI Benchmarking Hub
AIME (Vals)80.7%#47Vals AI
MGSM86.3%#59Vals AI

Knowledge

BenchmarkdefaultSource
AA-Omniscience-65.1#431Artificial Analysis
MMLU-Pro62.1%#128Vals AI
LegalBench40.0%#136Vals AI
CorpFin44.0%#113Vals AI
TaxEval60.3%#124Vals AI
MedQA82.4%#63Vals AI

Instruction following

BenchmarkdefaultSource
IFBench44.4%#209Artificial Analysis

Multimodal

BenchmarkdefaultSource
MMMU-Pro55.5%#195Artificial Analysis

Long context

BenchmarkdefaultSource
AA-LCR19.3%#370Artificial Analysis

Composite

BenchmarkdefaultSource
AA Intelligence Index8.6#364Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0007
Summarise a 30-page report12,000 / 600$0.0069
Code edit6,000 / 1,500$0.0053
Agentic coding session60,000 / 4,000$0.036
Structured extraction2,000 / 200$0.0013

See also

Data as of 10 Sept 2026. Compare with another model.