BenchLeader
Mistral AIAuto-detected

Magistral Medium 1.2

Magistral Medium 1.2 is a Mistral AI proprietary model, released 18 Sept 2025. Its best configuration ranks #403 of 372 on the BenchLeader Index at 45.2 ±6.0, in the lower half. It scores highest in long context (53) and lowest in knowledge (38). At $2.75 per million tokens blended it is pricier than most ranked models. Last measured 10 Sept 2026.

Blended price
$2.75/M
$2.00 in · $5.00 out
Output speed
First answer
Context
40k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index45
  2. Coding47
  3. Agents & tools45
  4. Maths40
  5. Knowledge38
  6. Instruction following45
  7. Multimodal44
  8. Long context53
  9. Composite44

Versions

Mistral AI has shipped 2 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
Magistral Medium 1.2this page18 Sept 202545.2#403
Magistral Mediummedium39.8#534

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkdefaultSource
GPQA Diamond (AA)not in index73.9%#255Artificial Analysis
Humanity's Last Exam (AA)not in index10.3%#258Artificial Analysis
GPQA Diamond (Vals)not in index62.4%#103Vals AI

Coding

BenchmarkdefaultSource
SciCode39.2%#116SciCode
LiveCodeBench74.9%#80Vals AI
IOI0.7%#56Vals AI

Agents & tools

Maths

BenchmarkdefaultSource
AIME (Vals)83.5%#42Vals AI
MGSM74.6%#68Vals AI

Knowledge

BenchmarkdefaultSource
AA-Omniscience-26.7#209Artificial Analysis
MMLU-Pro68.7%#121Vals AI
LegalBench54.8%#131Vals AI
CorpFin47.4%#105Vals AI
TaxEval61.9%#117Vals AI
MedQA89.5%#50Vals AI

Instruction following

BenchmarkdefaultSource
IFBench43.0%#225Artificial Analysis

Multimodal

BenchmarkdefaultSource
MMMU-Pro59.6%#183Artificial Analysis

Long context

BenchmarkdefaultSource
AA-LCR53.0%#247Artificial Analysis

Composite

BenchmarkdefaultSource
AA Intelligence Index11.8#280Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0023
Summarise a 30-page report12,000 / 600$0.027
Code edit6,000 / 1,500$0.019
Agentic coding session60,000 / 4,000$0.140
Structured extraction2,000 / 200$0.0050

See also

Data as of 10 Sept 2026. Compare with another model.