BenchLeader
Mistral AIOpen weightsAuto-detected

Devstral 2

Best configuration ranks #442 of 610 on the BenchLeader Index at 43.1 ±6.4. Last measured 8 Sept 2026. Released 9 Dec 2025.

Blended price
Output speed
126 tok/s
First answer
2.30 s
first token 2.30 s
Context
256k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index43
  2. Coding23
  3. Agents & tools50
  4. Knowledge42
  5. Instruction following41
  6. Long context42
  7. Composite41

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkdefaultSourceTrend
GPQA Diamond (AA)not in index59.4%#360Artificial Analysis
Humanity's Last Exam (AA)not in index3.6%#499Artificial Analysis

Coding

BenchmarkdefaultSourceTrend
LMArena WebDev1194#113LMArena
SciCode (AA)not in index32.8%#140Artificial Analysis

Agents & tools

BenchmarkdefaultSourceTrend
Terminal-Bench Hard18.9%#161Artificial Analysis
τ²-Bench Telecom (AA)not in index24.9%#285Artificial Analysis

Knowledge

BenchmarkdefaultSourceTrend
AA-Omniscience-46.7#293Artificial Analysis

Instruction following

BenchmarkdefaultSourceTrend
IFBench38.1%#269Artificial Analysis

Long context

BenchmarkdefaultSourceTrend
AA-LCR32.3%#312Artificial Analysis

Composite

BenchmarkdefaultSourceTrend
AA Intelligence Index9.4#330Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 3004.7 s
Summarise a 30-page report12,000 / 6007.1 s
Code edit6,000 / 1,50014.2 s
Agentic coding session60,000 / 4,00034.1 s
Structured extraction2,000 / 2003.9 s

See also

Data as of 9 Sept 2026. Compare with another model.