BenchLeader
OpenAIReasoning model

GPT-5 nano

Best configuration ranks #270 of 610 on the BenchLeader Index at 51.1 ±6.0. Last measured 7 Aug 2025stale: no new result in six months. Released 7 Aug 2025.

Blended price
$0.138/M
$0.050 in · $0.400 out
Output speed
176 tok/s
First answer
96 s
first token 2.19 s
Context
400k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index51
  2. Reasoning57
  3. Agents & tools48
  4. Knowledge50
  5. Instruction following66
  6. Multimodal45
  7. Long context48
  8. Composite45

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
minimal36.4#578168 tok/s0.88 s$0.0001Agents & tools 40 · Composite 38 · Instruction following 36 · Knowledge 34 · Long context 35 · Maths 32 · Multimodal 15 · Reasoning 33
low39.8#519176 tok/s96 s$0.0001Coding 34 · Maths 35 · Reasoning 36
medium46.5#367167 tok/s49 s$0.0001Agents & tools 35 · Coding 35 · Composite 45 · Instruction following 65 · Knowledge 52 · Long context 37 · Maths 62 · Multimodal 42 · Reasoning 40
high45.9#378176 tok/s96 s$0.0001Coding 48 · Human preference 51 · Knowledge 35 · Maths 48 · Multimodal 48 · Reasoning 42
defaultbest51.1#270176 tok/s96 s$0.0001Agents & tools 48 · Composite 45 · Instruction following 66 · Knowledge 50 · Long context 48 · Multimodal 45 · Reasoning 57

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkminimallowmediumhighdefaultSource
GPQA Diamond48.5%#20557.6%#18167.4%#16169.4%#158Epoch AI Benchmarking Hub
LMArena Hard Prompts1355#189LMArena
GPQA Diamond (AA)not in index42.8%#43867.0%#31167.6%#304Artificial Analysis
Humanity's Last Exam (AA)not in index4.0%#4678.7%#2779.5%#267Artificial Analysis
GPQA Diamond (Vals)not in index63.4%#102Vals AI
Kagi LLM Benchmark62.2%#46Kagi LLM Benchmark
ARC-AGI-11.5%#1844.0%#18120.7%#15716.7%#160ARC Prize
ARC-AGI-20.0%#1740.0%#1740.9%#1622.6%#137ARC Prize

Coding

BenchmarkminimallowmediumhighdefaultSource
WeirdML25.9%#12838.1%#111WeirdML
LMArena Coding1384#188LMArena
LiveCodeBench70.2%#86Vals AI
SWE-bench Verified (bash only)34.8%#36SWE-bench
SWE-bench Verified (any scaffold)not in index34.8%#54SWE-bench

Agents & tools

BenchmarkminimallowmediumhighdefaultSource
Terminal-Bench11.5%#6321.8%#53Terminal-Bench
Terminal-Bench Hard6.8%#24017.4%#17312.1%#202Artificial Analysis
τ²-Bench Telecom (AA)not in index25.7%#27930.4%#25136.5%#223Artificial Analysis
BFCL Overall51.5%#22Berkeley Function Calling Leaderboard

Maths

BenchmarkminimallowmediumhighdefaultSource
FrontierMath Tiers 1–31.8%#996.0%#9420.0%#74Epoch AI Benchmarking Hub
FrontierMath Tier 42.4%#53Epoch AI Benchmarking Hub
OTIS Mock AIME35.6%#18546.7%#16974.2%#11381.1%#98Epoch AI Benchmarking Hub
MATH Level 595.2%#1394.9%#15Epoch AI Benchmarking Hub
ProofBench12.0%#48Vals AI
AIME (Vals)81.2%#45Vals AI
MGSM89.3%#45Vals AI

Knowledge

BenchmarkminimallowmediumhighdefaultSource
SimpleQA Verified11.7%#71Epoch AI Benchmarking Hub
AA-Omniscience-64.1#415-25.8#200-28.7#210Artificial Analysis
MMLU-Pro76.1%#103Vals AI
LegalBench50.1%#134Vals AI
TaxEval67.4%#105Vals AI
MedQA93.3%#23Vals AI

Instruction following

BenchmarkminimallowmediumhighdefaultSource
IFBench32.5%#32665.9%#10067.5%#84Artificial Analysis

Human preference

BenchmarkminimallowmediumhighdefaultSource
LMArena Text1337#192LMArena

Multimodal

BenchmarkminimallowmediumhighdefaultSource
LMArena Vision1159#94LMArena
MMMU-Pro31.8%#23758.2%#18561.0%#174Artificial Analysis

Long context

BenchmarkminimallowmediumhighdefaultSource
Fiction.LiveBench 120k21.9%#33Fiction.live
AA-LCR20.0%#36443.7%#27445.0%#269Artificial Analysis

Composite

BenchmarkminimallowmediumhighdefaultSource
Epoch Capabilities Indexnot in index139.4#102Epoch AI Benchmarking Hub
AA Intelligence Index7.1#41412.5#26113.0#254Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
OpenAI Flex93 tok/s1.29 s$0.025$0.200400k
OpenAI90 tok/s2.48 s$0.050$0.400400k
Azure (EU)83 tok/s2.76 s$0.055$0.440400k
Azure55 tok/s1.91 s$0.050$0.400400k

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.121$0.242$0.363$0.484Jan 26Feb 26Mar 26Apr 26May 26Jun 26Jul 26Aug 26
input outputnow $0.055 in · $0.440 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.005 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0001$0.00011.6 min
Summarise a 30-page report12,000 / 600$0.0008$0.00041.7 min
Code edit6,000 / 1,500$0.0009$0.00071.7 min
Agentic coding session60,000 / 4,000$0.0046$0.00262.0 min
Structured extraction2,000 / 200$0.0002$0.00011.6 min

See also

Data as of 9 Sept 2026. Compare these configurations.