BenchLeader

GPT-4o mini

Best configuration ranks #589 of 610 on the BenchLeader Index at 35.8 ±4.3. Last measured 2 Sept 2026. Released 18 Jul 2024.

Blended price
$0.262/M
$0.150 in · $0.600 out
Output speed
147 tok/s
First answer
0.89 s
first token 1.17 s
Context
128k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index36
  2. Reasoning30
  3. Coding29
  4. Maths32
  5. Knowledge23
  6. Instruction following35
  7. Human preference48
  8. Multimodal33
  9. Composite37

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
high32.9#603147 tok/s0.89 s$0.0002Coding 14 · Knowledge 29 · Maths 35
defaultbest35.8#589147 tok/s0.89 s$0.0002Coding 29 · Composite 37 · Human preference 48 · Instruction following 35 · Knowledge 23 · Maths 32 · Multimodal 33 · Reasoning 30

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Coding

BenchmarkhighdefaultSource
WeirdML11.8%#142WeirdML
LMArena Coding1348#223LMArena
LiveCodeBench26.4%#132Vals AI
Aider Polyglot3.6%#45Aider polyglot leaderboard

Knowledge

BenchmarkhighdefaultSource
SimpleQA Verified8.3%#74Epoch AI Benchmarking Hub
MMLU-Pro62.7%#125Vals AI
CorpFin45.5%#110Vals AI
TaxEval60.5%#122Vals AI
MedQA72.4%#77Vals AI

Instruction following

BenchmarkhighdefaultSource
IFBench30.9%#341Artificial Analysis

Human preference

BenchmarkhighdefaultSource
LMArena Text1318#215LMArena

Multimodal

BenchmarkhighdefaultSource
LMArena Vision1066#113LMArena
MMMU-Pro41.5%#229Artificial Analysis
MMMU (validation)59.4%#33MMMU
MMMU-Pro (official)not in index37.6%#23MMMU

Composite

BenchmarkhighdefaultSource
Epoch Capabilities Indexnot in index126.6#147Epoch AI Benchmarking Hub
AA Intelligence Index6.7#434Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
OpenAI47 tok/s0.59 s$0.150$0.600128k
Azure21 tok/s1.75 s$0.150$0.600128k

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.182$0.363$0.545$0.726Jan 26Feb 26Mar 26Apr 26May 26Jun 26Jul 26Aug 26
input outputnow $0.165 in · $0.660 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.075 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0002$0.00022.9 s
Summarise a 30-page report12,000 / 600$0.0022$0.00155.0 s
Code edit6,000 / 1,500$0.0018$0.001511.1 s
Agentic coding session60,000 / 4,000$0.011$0.008028.1 s
Structured extraction2,000 / 200$0.0004$0.00032.2 s

See also

Data as of 9 Sept 2026. Compare these configurations.