BenchLeader
OpenAIReasoning modelAuto-detected

GPT-5.4 mini

Best configuration ranks #191 of 610 on the BenchLeader Index at 54.8 ±8.3. Last measured 17 Mar 2026. Released 17 Mar 2026.

Blended price
$1.69/M
$0.750 in · $4.50 out
Output speed
218 tok/s
First answer
139 s
first token 1.31 s
Context
400k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index55
  2. Reasoning36
  3. Coding42
  4. Agents & tools65
  5. Knowledge55
  6. Instruction following71
  7. Multimodal58
  8. Long context65
  9. Composite60

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning44.9#402145 tok/s0.71 s$0.0017Agents & tools 49 · Coding 42 · Composite 43 · Instruction following 41 · Knowledge 43 · Long context 44 · Maths 36 · Multimodal 45 · Reasoning 48
low218 tok/s139 s$0.0017Maths 41 · Reasoning 32
medium53.9#211173 tok/s6.27 s$0.0017Agents & tools 63 · Composite 54 · Instruction following 64 · Knowledge 54 · Long context 60 · Multimodal 56 · Reasoning 38
high52.9#236218 tok/s139 s$0.0017Agents & tools 44 · Coding 53 · Human preference 64 · Knowledge 41 · Maths 62 · Multimodal 59 · Reasoning 53
xhigh46.0#372218 tok/s139 s$0.0017Agents & tools 33 · Coding 54 · Composite 25 · Knowledge 55 · Maths 53 · Reasoning 53
defaultbest54.8#191218 tok/s139 s$0.0017Agents & tools 65 · Coding 42 · Composite 60 · Instruction following 71 · Knowledge 55 · Long context 65 · Multimodal 58 · Reasoning 36

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
GPQA Diamond64.1%#17083.6%#8886.9%#65Epoch AI Benchmarking Hub
LMArena Hard Prompts1470#70LMArena
LiveBench Reasoningnot in index71.3%#49LiveBench
GPQA Diamond (AA)not in index60.6%#35282.3%#16887.5%#96Artificial Analysis
Humanity's Last Exam (AA)not in index5.9%#33818.6%#17128.1%#119Artificial Analysis
GPQA Diamond (Vals)not in index83.1%#57Vals AI
Kagi LLM Benchmark37.9%#108Kagi LLM Benchmark
ARC-AGI-113.0%#16540.8%#13158.0%#11163.7%#104ARC Prize
ARC-AGI-21.1%#1614.4%#12513.2%#10118.9%#94ARC Prize

Coding

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
SciCode49.9%#62SciCode
WeirdML37.9%#11460.3%#47WeirdML
FrontierCode27.0%#19Cognition
LMArena Coding1496#70LMArena
LMArena WebDev1397#72LMArena
LiveBench Codingnot in index71.6%#45LiveBench
SciCode (AA)not in index52.1%#55Artificial Analysis
LiveCodeBench81.5%#59Vals AI
IOI6.4%#41Vals AI
SWE-bench (Vals)not in index73.0%#51Vals AI

Agents & tools

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
APEX-Agents24.6%#30Mercor
LiveBench Agentic Codingnot in index41.7%#46LiveBench
Terminal-Bench Hard18.2%#16634.1%#8752.3%#19Artificial Analysis
τ²-Bench Telecom (AA)not in index23.4%#29136.5%#22383.3%#106Artificial Analysis
Terminal-Bench 2.1 (Vals)54.7%#40Vals AI
MCP Atlas56.7%#27Scale AI SEAL

Maths

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
FrontierMath Tiers 1–317.2%#8424.6%#6651.2%#38Epoch AI Benchmarking Hub
FrontierMath Tier 49.8%#50Epoch AI Benchmarking Hub
OTIS Mock AIME26.7%#19387.2%#6988.9%#61Epoch AI Benchmarking Hub
ProofBench21.0%#34Vals AI
LiveBench Mathematicsnot in index78.5%#48LiveBench
AIME (Vals)95.6%#8Vals AI

Knowledge

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
SimpleQA Verified29.4%#58Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index70.8%#36LiveBench
AA-Omniscience-45.4#286-20.6#184-18.9#178Artificial Analysis
MMLU-Pro84.5%#61Vals AI
CorpFin60.9%#62Vals AI
TaxEval71.2%#78Vals AI

Instruction following

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
LiveBench Languagenot in index71.0%#48LiveBench
IFBench38.8%#26464.8%#10473.3%#42Artificial Analysis

Human preference

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
LMArena Text1448#70LMArena

Multimodal

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
LMArena Vision1246#53LMArena
MMMU-Pro60.5%#17571.2%#10973.3%#96Artificial Analysis

Long context

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
AA-LCR37.0%#29967.0%#18577.0%#95Artificial Analysis

Composite

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
Epoch Capabilities Indexnot in index149.0#50Epoch AI Benchmarking Hub
LiveBench66.4%#49LiveBench
AA Intelligence Index11.1#29319.8#17624.6#123Artificial Analysis
Vals Indexnot in index39.6#36Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
OpenAI Fast78 tok/s1.47 s$1.50$9.00400k
OpenAI59 tok/s1.16 s$0.750$4.50400k
Azure (US)43 tok/s1.68 s$0.825$4.95400k
Azure39 tok/s1.08 s$0.750$4.50400k

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.075 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0017$0.00142.3 min
Summarise a 30-page report12,000 / 600$0.012$0.00562.4 min
Code edit6,000 / 1,500$0.011$0.00822.4 min
Agentic coding session60,000 / 4,000$0.063$0.0332.6 min
Structured extraction2,000 / 200$0.0024$0.00142.3 min

See also

Data as of 9 Sept 2026. Compare these configurations.