BenchLeader

GPT-4.1 nano

Best configuration ranks #554 of 610 on the BenchLeader Index at 38.1 ±4.0. Last measured 2 Sept 2026. Released 14 Apr 2025.

Blended price
$0.175/M
$0.100 in · $0.400 out
Output speed
114 tok/s
First answer
0.76 s
first token 1.27 s
Context
1.0M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index38
  2. Reasoning35
  3. Coding31
  4. Agents & tools43
  5. Maths46
  6. Knowledge29
  7. Instruction following36
  8. Human preference49
  9. Multimodal28
  10. Long context30
  11. Composite39

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
high31.7#607114 tok/s0.76 s$0.0002Coding 27 · Knowledge 27 · Maths 21
defaultbest38.1#554114 tok/s0.76 s$0.0002Agents & tools 43 · Coding 31 · Composite 39 · Human preference 49 · Instruction following 36 · Knowledge 29 · Long context 30 · Maths 46 · Multimodal 28 · Reasoning 35

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Coding

BenchmarkhighdefaultSource
SciCode25.9%#148SciCode
WeirdML19.0%#137WeirdML
LMArena Coding1374#196LMArena
LiveCodeBench42.7%#120Vals AI
Aider Polyglot8.9%#43Aider polyglot leaderboard

Agents & tools

Maths

BenchmarkhighdefaultSource
OTIS Mock AIME28.9%#191Epoch AI Benchmarking Hub
MATH Level 570.0%#38Epoch AI Benchmarking Hub
AIME (Vals)26.5%#74Vals AI
MGSM69.3%#69Vals AI

Knowledge

BenchmarkhighdefaultSource
SimpleQA Verified6.0%#75Epoch AI Benchmarking Hub
AA-Omniscience-57.6#369Artificial Analysis
MMLU-Pro63.5%#124Vals AI
LegalBench61.1%#128Vals AI
CorpFin42.1%#115Vals AI
TaxEval60.8%#121Vals AI
MedQA68.2%#79Vals AI

Instruction following

BenchmarkhighdefaultSource
IFBench32.0%#328Artificial Analysis

Human preference

BenchmarkhighdefaultSource
LMArena Text1322#210LMArena

Multimodal

BenchmarkhighdefaultSource
LMArena Vision1063#114LMArena
MMMU-Pro40.1%#230Artificial Analysis
VISTA26.6%#51Scale AI SEAL

Long context

BenchmarkhighdefaultSource
Fiction.LiveBench 120k18.8%#35Fiction.live
AA-LCR20.3%#360Artificial Analysis

Composite

BenchmarkhighdefaultSource
Epoch Capabilities Indexnot in index129.6#134Epoch AI Benchmarking Hub
AA Intelligence Index7.8#378Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
OpenAI62 tok/s1.23 s$0.100$0.4001.0M
Azure14 tok/s1.31 s$0.100$0.4001.0M

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.121$0.242$0.363$0.484Jan 26Feb 26Mar 26Apr 26May 26Jun 26Jul 26Aug 26
input outputnow $0.110 in · $0.440 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.025 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0002$0.00013.4 s
Summarise a 30-page report12,000 / 600$0.0014$0.00086.0 s
Code edit6,000 / 1,500$0.0012$0.000913.9 s
Agentic coding session60,000 / 4,000$0.0076$0.004235.9 s
Structured extraction2,000 / 200$0.0003$0.00022.5 s

See also

Data as of 9 Sept 2026. Compare these configurations.