BenchLeader
OpenAIReasoning model

o3-mini

Best configuration ranks #307 of 610 on the BenchLeader Index at 49.5 ±5.2. Last measured 2 Sept 2026. Released 31 Jan 2025.

Blended price
$1.93/M
$1.10 in · $4.40 out
Output speed
213 tok/s
First answer
5.59 s
first token 5.49 s
Context
200k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index50
  2. Reasoning53
  3. Coding58
  4. Agents & tools40
  5. Human preference52
  6. Composite45

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
low40.4#507213 tok/s5.59 s$0.0018Coding 34 · Maths 34 · Reasoning 40
medium45.6#386213 tok/s5.59 s$0.0018Agents & tools 46 · Long context 41 · Maths 46 · Reasoning 42
high47.8#344213 tok/s21 s$0.0018Agents & tools 39 · Coding 47 · Composite 43 · Human preference 54 · Instruction following 66 · Knowledge 42 · Long context 47 · Maths 50 · Reasoning 42
defaultbest49.5#307213 tok/s5.59 s$0.0018Agents & tools 40 · Coding 58 · Composite 45 · Human preference 52 · Reasoning 53

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarklowmediumhighdefaultSource
GPQA Diamond68.2%#16074.3%#14077.0%#127Epoch AI Benchmarking Hub
SimpleBench22.8%#75SimpleBench
LMArena Hard Prompts1401#1521371#174LMArena
GPQA Diamond (AA)not in index77.3%#21474.8%#245Artificial Analysis
Humanity's Last Exam (AA)not in index12.0%#2207.9%#285Artificial Analysis
GPQA Diamond (Vals)not in index75.5%#77Vals AI
ARC-AGI-114.5%#16222.3%#15334.5%#138ARC Prize
ARC-AGI-20.0%#1742.1%#1413.0%#135ARC Prize

Coding

BenchmarklowmediumhighdefaultSource
SciCode39.8%#112SciCode
WeirdML43.7%#86WeirdML
GSO-Bench1.3%#291.3%#29GSO-Bench
LMArena Coding1435#1461417#159LMArena
SciCode (AA)not in index42.8%#103Artificial Analysis
LiveCodeBench71.5%#82Vals AI
Aider Polyglot60.4%#14Aider polyglot leaderboard
SWE-bench Verified (any scaffold)not in index42.4%#46SWE-bench

Agents & tools

BenchmarklowmediumhighdefaultSource
Cybench22.5%#10Cybench
Terminal-Bench Hard6.1%#2576.8%#240Artificial Analysis
τ²-Bench Telecom (AA)not in index31.3%#24728.6%#259Artificial Analysis

Maths

BenchmarklowmediumhighdefaultSource
FrontierMath Tiers 1–33.9%#9810.5%#8918.6%#79Epoch AI Benchmarking Hub
FrontierMath Tier 40.0%#57Epoch AI Benchmarking Hub
OTIS Mock AIME44.4%#17563.9%#14076.9%#110Epoch AI Benchmarking Hub
MATH Level 595.2%#1496.5%#10Epoch AI Benchmarking Hub
AIME (Vals)86.5%#32Vals AI
MGSM91.3%#27Vals AI

Knowledge

BenchmarklowmediumhighdefaultSource
SimpleQA Verified15.3%#67Epoch AI Benchmarking Hub
AA-Omniscience-42.6#269Artificial Analysis
MMLU-Pro78.7%#97Vals AI
LegalBench71.5%#110Vals AI
CorpFin45.3%#111Vals AI
TaxEval69.4%#95Vals AI
MedQA94.8%#14Vals AI

Instruction following

BenchmarklowmediumhighdefaultSource
IFBench67.1%#88Artificial Analysis

Human preference

BenchmarklowmediumhighdefaultSource
LMArena Text1363#1671348#181LMArena

Long context

BenchmarklowmediumhighdefaultSource
Fiction.LiveBench 120k43.8%#24Fiction.live
AA-LCR43.0%#277Artificial Analysis

Composite

BenchmarklowmediumhighdefaultSource
Epoch Capabilities Indexnot in index140.4#96Epoch AI Benchmarking Hub
AA Intelligence Index11.0#29812.5#263Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
OpenAI109 tok/s5.49 s$1.10$4.40200k

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.550 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0018$0.00167.0 s
Summarise a 30-page report12,000 / 600$0.016$0.0118.4 s
Code edit6,000 / 1,500$0.013$0.01112.6 s
Agentic coding session60,000 / 4,000$0.084$0.05924.4 s
Structured extraction2,000 / 200$0.0031$0.00236.5 s

See also

Data as of 9 Sept 2026. Compare these configurations.