BenchLeader
Mistral AIOpen weightsReasoning model

Magistral Medium

Best configuration ranks #421 of 610 on the BenchLeader Index at 44.1 ±5.2 (medium reasoning effort). Last measured 2 Sept 2026. Released 17 Mar 2025.

Blended price
$2.75/M
$2.00 in · $5.00 out
Output speed
First answer
Context
40k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index44
  2. Reasoning32
  3. Coding47
  4. Agents & tools45
  5. Maths40
  6. Knowledge38
  7. Instruction following45
  8. Human preference47
  9. Multimodal44
  10. Long context52
  11. Composite44

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Coding

BenchmarkmediumSource
SciCode39.2%#114SciCode
LMArena Coding1387#184LMArena
LiveCodeBench74.9%#79Vals AI
IOI0.7%#57Vals AI

Agents & tools

Maths

BenchmarkmediumSource
AIME (Vals)83.5%#42Vals AI
MGSM74.6%#67Vals AI

Knowledge

BenchmarkmediumSource
AA-Omniscience-26.7#205Artificial Analysis
MMLU-Pro68.7%#120Vals AI
LegalBench54.8%#130Vals AI
CorpFin47.4%#105Vals AI
TaxEval61.9%#116Vals AI
MedQA89.5%#50Vals AI

Instruction following

BenchmarkmediumSource
IFBench43.0%#225Artificial Analysis

Human preference

BenchmarkmediumSource
LMArena Text1304#229LMArena

Multimodal

BenchmarkmediumSource
MMMU-Pro59.6%#180Artificial Analysis

Long context

BenchmarkmediumSource
AA-LCR53.0%#243Artificial Analysis

Composite

BenchmarkmediumSource
AA Intelligence Index11.8#276Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0023
Summarise a 30-page report12,000 / 600$0.027
Code edit6,000 / 1,500$0.019
Agentic coding session60,000 / 4,000$0.140
Structured extraction2,000 / 200$0.0050

See also

Data as of 9 Sept 2026. Compare with another model.