BenchLeader
Mistral AIOpen weightsReasoning modelNewReleased 6 Oct 2026Compare with another model

Mistral Large 4

Reasoning effort

Mistral Large 4 is a Mistral AI open-weights reasoning model, released 6 Oct 2026. Its best configuration ranks #159 of 760 on the BenchLeader Index at 57.8 ±5.7, in the upper half. The ± is the point: 274 other configurations score within that range, so they and this one cannot be told apart on quality alone — price and speed are what separate them. It scores highest in composite (73) and lowest in agents & tools (33). At $2.06 per million tokens blended it is pricier than most ranked models. Output speed of 55 tokens per second puts it slower than most, with a first token in 1.8 s. It has been measured at 2 reasoning-effort settings; this summary describes the best-scoring one, and the tabs above switch between them. Last measured 8 Oct 2026.

Blended price
$2.06/M
$1.36 in · $4.18 out
Output speed
55 tok/s
OpenRouter traffic, 7-day median; not yet measured by Artificial Analysis
First answer
1.78 s
Context
1.0M
Full answer
–
Cost per run
$1.13
one full Intelligence Index run
Released
6 Oct 2026
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index58
  2. Reasoning60
  3. Coding60
  4. Agents & tools33
  5. Maths67
  6. Knowledge59
  7. Human preference61
  8. Multimodal60
  9. Long context66
  10. Composite73

Versions

Mistral AI has shipped 4 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
NewMistral Large 4this page6 Oct 202657.8#159
Mistral Large 31 Nov 202445.5#489
Mistral Large 21 Nov 202440.6#632
Mistral Large 240226 Feb 202441.0#622

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Not stated” collects results from publishers that did not say which setting they used; for a reasoning model that is usually its thinking mode, but we do not assume it. Pick a setting here or at the top of the page to see its price, speed and category scores.

EffortIndexRankSpeedFirst answerChat reply costCategories
high––55 tok/s1.78 s$0.0018Composite 41 · Knowledge 33
not statedbest57.8#15955 tok/s1.78 s$0.0018Agents & tools 33 · Coding 60 · Composite 73 · Human preference 61 · Knowledge 59 · Long context 66 · Maths 67 · Multimodal 60 · Reasoning 60

How its index has moved

6 Oct 2026 to 9 Oct 2026
525864

The index is recomputed from scratch every day, so a line moves when a new benchmark result lands, when a publisher revises a score, or when the models it is normalised against change. Early movement usually means the score is still settling.

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkhighnot statedSourceTrend
LMArena Hard Promptseffort not stated–1453#104LMArena
LiveBench Reasoningnot in index83.9%#41–LiveBench
Humanity's Last Exam (AA)not in indexeffort not stated–35.0%#113Artificial Analysis
CritPteffort not stated–10.6%#108Artificial Analysis

Coding

Benchmarkhighnot statedSourceTrend
LMArena Codingeffort not stated–1492#86LMArena
LMArena WebDeveffort not stated–1541#41LMArena
LiveBench Codingnot in index77.2%#38–LiveBench
SciCode (AA)not in indexeffort not stated–54.2%#61Artificial Analysis
Terminal-Bench 4.0 (AA)not in indexeffort not stated–26.8%#45Artificial Analysis

Agents & tools

Benchmarkhighnot statedSourceTrend
LMArena Agenteffort not stated–-6.6#41LMArena
LiveBench Agentic Codingnot in index57.2%#21–LiveBench
GDPval-AA v2.1not in indexeffort not stated–46.2%#64Artificial Analysis
AutomationBenchnot in indexeffort not stated–59.9%#22Artificial Analysis
GDP.pdfnot in indexeffort not stated–18.6%#31Artificial Analysis
Harvey LABnot in indexeffort not stated–3.1%#10Harvey
AA-Briefcase v1.1not in indexeffort not stated–1393#37Artificial Analysis

Maths

Benchmarkhighnot statedSourceTrend
LiveBench Mathematicsnot in index93.6%#20–LiveBench
LMArena Mathseffort not stated–1487#30LMArena

Knowledge

Benchmarkhighnot statedSourceTrend
SimpleQA Verified20.0%#70–Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index76.5%#31–LiveBench
AA-Omniscienceeffort not stated–-5.3#157Artificial Analysis
AA-Omniscience: accuracynot in indexeffort not stated–25.8%#211Artificial Analysis
AA-Omniscience: non-hallucinationnot in indexeffort not stated–58.0%#79Artificial Analysis

Instruction following

Benchmarkhighnot statedSourceTrend
LiveBench Languagenot in index49.6%#66–LiveBench
LiveBench Instruction Followingnot in index64.8%#44–LiveBench
LMArena Instruction Followingnot in indexeffort not stated–1426#97LMArena

Human preference

Benchmarkhighnot statedSourceTrend
LMArena Texteffort not stated–1429#109LMArena
LMArena Creative Writingnot in indexeffort not stated–1367#145LMArena
LMArena Multi-turnnot in indexeffort not stated–1425#117LMArena
LMArena Longer Queriesnot in indexeffort not stated–1435#111LMArena

Multimodal

Benchmarkhighnot statedSourceTrend
MMMU-Proeffort not stated–76.4%#84Artificial Analysis

Long context

Benchmarkhighnot statedSourceTrend
AA-LCReffort not stated–81.3%#52Artificial Analysis

Composite

Benchmarkhighnot statedSourceTrend
LiveBench71.8%#51–LiveBench
AA Intelligence Index v4.3.2effort not stated–38.4#70Artificial Analysis

Safety & honesty

Benchmarkhighnot statedSourceTrend
FORTRESSnot in indexeffort not stated–23.1%#34Scale AI SEAL

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Mistral79 tok/s0.79 s$0.680$2.091.0M–
Mistral70 tok/s1.60 s$0.680$2.091.0M–

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.070 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0018$0.00147.2 s
Summarise a 30-page report12,000 / 600$0.019$0.007212.7 s
Code edit6,000 / 1,500$0.014$0.008629.1 s
Agentic coding session60,000 / 4,000$0.098$0.0401.2 min
Structured extraction2,000 / 200$0.0036$0.00165.4 s

See also

Data as of 11 Oct 2026. Compare these configurations.

Cite as: BenchLeader, “Mistral Large 4: benchmarks, pricing, speed and rank”, https://www.benchleader.com/models/mistral-large-4, data as of 11 Oct 2026.