BenchLeader
OpenAIReasoning model

GPT-5 mini

Best configuration ranks #172 of 610 on the BenchLeader Index at 55.8 ±5.0. Last measured 17 Feb 2026stale: no new result in six months. Released 7 Aug 2025.

Blended price
$0.688/M
$0.250 in · $2.00 out
Output speed
103 tok/s
First answer
74 s
first token 6.81 s
Context
400k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index56
  2. Reasoning56
  3. Coding48
  4. Agents & tools56
  5. Knowledge50
  6. Instruction following73
  7. Multimodal59
  8. Long context63
  9. Composite51

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
minimal43.9#428102 tok/s0.94 s$0.0007Agents & tools 46 · Composite 42 · Instruction following 47 · Knowledge 39 · Long context 45 · Maths 36 · Multimodal 43 · Reasoning 40
low103 tok/s74 s$0.0007Maths 37 · Reasoning 34
medium54.5#19695 tok/s13 s$0.0007Agents & tools 48 · Coding 49 · Composite 55 · Instruction following 69 · Knowledge 59 · Long context 58 · Maths 64 · Multimodal 53 · Reasoning 43
high52.3#250103 tok/s74 s$0.0007Coding 55 · Human preference 57 · Knowledge 50 · Maths 52 · Multimodal 53 · Reasoning 48
thinking103 tok/s74 s$0.0007Instruction following 54
defaultbest55.8#172103 tok/s74 s$0.0007Agents & tools 56 · Coding 48 · Composite 51 · Instruction following 73 · Knowledge 50 · Long context 63 · Multimodal 59 · Reasoning 56

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkminimallowmediumhighthinkingdefaultSource
GPQA Diamond71.7%#15071.7%#15275.0%#138Epoch AI Benchmarking Hub
Humanity's Last Exam19.4%#18Scale AI / CAIS
LMArena Hard Prompts1403#150LMArena
GPQA Diamond (AA)not in index68.7%#29780.3%#18782.8%#163Artificial Analysis
Humanity's Last Exam (AA)not in index5.0%#37415.9%#18621.5%#154Artificial Analysis
GPQA Diamond (Vals)not in index80.3%#65Vals AI
Kagi LLM Benchmark70.3%#26Kagi LLM Benchmark
ARC-AGI-15.3%#17726.3%#14937.3%#13654.3%#118ARC Prize
ARC-AGI-21.7%#1480.8%#1644.0%#1274.4%#125ARC Prize

Coding

BenchmarkminimallowmediumhighthinkingdefaultSource
SWE-bench Verified (Epoch)64.7%#27Epoch AI Benchmarking Hub
SciCode39.2%#114SciCode
WeirdML52.7%#60WeirdML
LMArena Coding1431#152LMArena
SciCode (AA)not in index39.0%#119Artificial Analysis
LiveCodeBench86.6%#20Vals AI
IOI6.8%#38Vals AI
SWE-bench (Vals)not in index60.8%#73Vals AI
SWE-bench Verified (bash only)59.8%#2456.2%#27SWE-bench
SWE-bench Verified (any scaffold)not in index59.8%#32SWE-bench

Agents & tools

BenchmarkminimallowmediumhighthinkingdefaultSource
Terminal-Bench31.9%#4534.8%#42Terminal-Bench
Terminal-Bench Hard14.4%#18928.8%#11533.3%#90Artificial Analysis
τ²-Bench Telecom (AA)not in index31.9%#24371.0%#14868.4%#156Artificial Analysis
BFCL Overall55.5%#15Berkeley Function Calling Leaderboard

Maths

BenchmarkminimallowmediumhighthinkingdefaultSource
FrontierMath Tiers 1–36.0%#9418.3%#8046.7%#40Epoch AI Benchmarking Hub
FrontierMath Tier 412.2%#47Epoch AI Benchmarking Hub
OTIS Mock AIME55.6%#15778.3%#10686.7%#71Epoch AI Benchmarking Hub
MATH Level 596.8%#897.8%#3Epoch AI Benchmarking Hub
ProofBench9.0%#50Vals AI
AIME (Vals)91.5%#24Vals AI
MGSM92.6%#15Vals AI
MathArena Apex1.0%#35MathArena

Knowledge

BenchmarkminimallowmediumhighthinkingdefaultSource
SimpleQA Verified21.6%#61Epoch AI Benchmarking Hub
AA-Omniscience-53.0#341-11.0#149-17.3#172Artificial Analysis
MMLU-Pro82.2%#75Vals AI
LegalBench81.8%#73Vals AI
CorpFin60.2%#70Vals AI
TaxEval75.2%#21Vals AI
MedQA96.1%#6Vals AI
MultiNRC23.9%#30Scale AI SEAL

Instruction following

BenchmarkminimallowmediumhighthinkingdefaultSource
IFBench45.6%#19971.2%#6075.4%#31Artificial Analysis
MultiChallenge59.0%#11Scale AI SEAL

Human preference

BenchmarkminimallowmediumhighthinkingdefaultSource
LMArena Text1389#143LMArena

Multimodal

BenchmarkminimallowmediumhighthinkingdefaultSource
LMArena Vision1202#72LMArena
MMMU-Pro58.4%#18468.8%#12870.1%#116Artificial Analysis
VISTA50.4%#9Scale AI SEAL

Long context

BenchmarkminimallowmediumhighthinkingdefaultSource
Fiction.LiveBench 120k62.5%#13Fiction.live
AA-LCR39.3%#29172.3%#13772.3%#137Artificial Analysis

Composite

BenchmarkminimallowmediumhighthinkingdefaultSource
Epoch Capabilities Indexnot in index145.5#69Epoch AI Benchmarking Hub
AA Intelligence Index9.9#31520.6#16917.4#199Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
OpenAI78 tok/s4.02 s$0.250$2.00400k
Azure51 tok/s6.81 s$0.250$2.00400k
OpenAI Flex33 tok/s15 s$0.125$1.00400k

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.605$1.21$1.82$2.42Jan 26Feb 26Mar 26Apr 26May 26Jun 26Jul 26Aug 26
input outputnow $0.275 in · $2.20 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.025 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0007$0.00061.3 min
Summarise a 30-page report12,000 / 600$0.0042$0.00221.3 min
Code edit6,000 / 1,500$0.0045$0.00351.5 min
Agentic coding session60,000 / 4,000$0.023$0.0131.9 min
Structured extraction2,000 / 200$0.0009$0.00061.3 min

See also

Data as of 9 Sept 2026. Compare these configurations.