BenchLeader
Mistral AIOpen weightsAuto-detected

Mistral Small 4

Best configuration ranks #357 of 610 on the BenchLeader Index at 47.1 ±3.3. Released 16 Mar 2026.

Blended price
$0.262/M
$0.150 in · $0.600 out
Output speed
168 tok/s
First answer
13 s
Context
256k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index47
  2. Reasoning39
  3. Agents & tools49
  4. Knowledge50
  5. Instruction following49
  6. Multimodal41
  7. Long context51
  8. Composite43

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning42.2#466153 tok/s0.73 s$0.0002Agents & tools 43 · Composite 40 · Instruction following 36 · Knowledge 41 · Long context 39 · Multimodal 30
defaultbest47.1#357168 tok/s13 s$0.0002Agents & tools 49 · Composite 43 · Instruction following 49 · Knowledge 50 · Long context 51 · Multimodal 41 · Reasoning 39

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoningdefaultSource
GPQA Diamond (AA)not in index57.1%#37676.9%#219Artificial Analysis
Humanity's Last Exam (AA)not in index3.8%#4799.9%#260Artificial Analysis
Kagi LLM Benchmark40.5%#102Kagi LLM Benchmark

Coding

Benchmarkno reasoningdefaultSource
SciCode (AA)not in index38.8%#122Artificial Analysis

Agents & tools

Benchmarkno reasoningdefaultSource
Terminal-Bench Hard10.6%#21317.4%#173Artificial Analysis
τ²-Bench Telecom (AA)not in index18.4%#33241.2%#214Artificial Analysis

Knowledge

Benchmarkno reasoningdefaultSource
AA-Omniscience-48.5#303-30.4#217Artificial Analysis

Instruction following

Benchmarkno reasoningdefaultSource
IFBench32.8%#32148.2%#182Artificial Analysis

Multimodal

Benchmarkno reasoningdefaultSource
MMMU-Pro46.5%#21656.8%#188Artificial Analysis

Long context

Benchmarkno reasoningdefaultSource
AA-LCR28.3%#32649.7%#256Artificial Analysis

Composite

Benchmarkno reasoningdefaultSource
AA Intelligence Index9.0#34211.4#286Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.000214.4 s
Summarise a 30-page report12,000 / 600$0.002216.2 s
Code edit6,000 / 1,500$0.001821.5 s
Agentic coding session60,000 / 4,000$0.01136.4 s
Structured extraction2,000 / 200$0.000413.8 s

See also

Data as of 9 Sept 2026. Compare these configurations.