BenchLeader
OpenAIReasoning modelAuto-detected

GPT-5.4

Best configuration ranks #46 of 610 on the BenchLeader Index at 63.4 ±6.2. Last measured 8 Sept 2026. Released 5 Mar 2026.

Blended price
$5.63/M
$2.50 in · $15.00 out
Output speed
130 tok/s
First answer
126 s
first token 3.21 s
Context
1.1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index63
  2. Reasoning62
  3. Coding52
  4. Agents & tools66
  5. Knowledge67
  6. Instruction following72
  7. Human preference65
  8. Multimodal64
  9. Long context68
  10. Composite78

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning53.7#216101 tok/s0.83 s$0.0055Agents & tools 67 · Coding 55 · Composite 52 · Instruction following 50 · Knowledge 57 · Long context 55 · Maths 49 · Multimodal 55 · Reasoning 55
low60.4#88106 tok/s1.68 s$0.0055Agents & tools 71 · Composite 64 · Instruction following 65 · Knowledge 66 · Long context 65 · Maths 61 · Multimodal 63 · Reasoning 54
medium56.0#168130 tok/s126 s$0.0061Coding 51 · Maths 66 · Reasoning 62
high58.8#115130 tok/s126 s$0.0061Agents & tools 51 · Coding 58 · Human preference 68 · Knowledge 53 · Maths 67 · Multimodal 66 · Reasoning 62
xhigh61.3#67130 tok/s126 s$0.0061Agents & tools 54 · Coding 68 · Composite 60 · Knowledge 59 · Maths 67 · Reasoning 68
thinking130 tok/s126 s$0.0061Knowledge 70 · Multimodal 65
defaultbest63.4#46130 tok/s126 s$0.0055Agents & tools 66 · Coding 52 · Composite 78 · Human preference 65 · Instruction following 72 · Knowledge 67 · Long context 68 · Multimodal 64 · Reasoning 62

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowmediumhighxhighthinkingdefaultSourceTrend
GPQA Diamond74.8%#13984.8%#8188.9%#4589.9%#3993.3%#15Epoch AI Benchmarking Hub
Humanity's Last Exam36.2%#6Scale AI / CAIS
LMArena Hard Prompts1498#271488#44LMArena
LiveBench Reasoningnot in index88.1%#17LiveBench
GPQA Diamond (AA)not in index74.8%#24587.1%#10292.0%#41Artificial Analysis
Humanity's Last Exam (AA)not in index11.3%#23230.8%#9743.7%#31Artificial Analysis
GPQA Diamond (Vals)not in index91.7%#18Vals AI
Kagi LLM Benchmark63.8%#40Kagi LLM Benchmark
ARC-AGI-168.2%#9886.2%#6892.7%#3493.7%#31ARC Prize
ARC-AGI-229.2%#9055.4%#6567.5%#4174.0%#32ARC Prize
ARC-AGI-30.2%#30ARC Prize

Coding

Benchmarkno reasoninglowmediumhighxhighthinkingdefaultSourceTrend
SWE-bench Verified (Epoch)76.9%#8Epoch AI Benchmarking Hub
SciCode56.6%#12SciCode
WeirdML57.4%#5077.7%#21WeirdML
GSO-Bench25.5%#1131.4%#8GSO-Bench
LMArena Coding1520#271514#41LMArena
LMArena WebDev1443#561463#501386#77LMArena
LiveBench Codingnot in index77.5%#26LiveBench
LiveCodeBench84.1%#39Vals AI
IOI67.8%#6Vals AI
SWE-bench (Vals)not in index78.2%#30Vals AI
SWE-Bench Pro59.1%#2Scale AI SEAL

Agents & tools

Benchmarkno reasoninglowmediumhighxhighthinkingdefaultSourceTrend
Terminal-Bench81.8%#2Terminal-Bench
APEX-Agents36.0%#1634.9%#18Mercor
LMArena Agent1.6#21LMArena
LiveBench Agentic Codingnot in index53.8%#26LiveBench
Terminal-Bench Hard37.9%#6143.2%#3957.6%#11Artificial Analysis
τ²-Bench Telecom (AA)not in index36.0%#22874.6%#13587.1%#79Artificial Analysis
MCP Atlas70.6%#17Scale AI SEAL
HiL-Bench9.7%#15Scale AI SEAL

Maths

Benchmarkno reasoninglowmediumhighxhighthinkingdefaultSourceTrend
FrontierMath Tiers 1–378.6%#12Epoch AI Benchmarking Hub
FrontierMath Tier 449.0%#19Epoch AI Benchmarking Hub
OTIS Mock AIME57.8%#15084.4%#8395.6%#2997.8%#1795.3%#37Epoch AI Benchmarking Hub
ProofBench56.0%#13Vals AI
LiveBench Mathematicsnot in index94.2%#11LiveBench
AIME (Vals)96.7%#5Vals AI
AIME 202699.2%#3MathArena
HMMT February 202697.7%#2MathArena
MathArena Apex54.2%#6MathArena

Knowledge

Benchmarkno reasoninglowmediumhighxhighthinkingdefaultSourceTrend
SimpleQA Verified45.1%#32Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index79.3%#9LiveBench
AA-Omniscience-15.7#1654.8#735.8#69Artificial Analysis
MMLU-Pro87.5%#27Vals AI
LegalBench86.0%#13Vals AI
CorpFin65.3%#32Vals AI
TaxEval74.0%#42Vals AI
MedQA96.1%#5Vals AI
PRBench Finance45.6%#17Scale AI SEAL
PRBench Legal44.4%#15Scale AI SEAL
MultiNRC58.3%#7Scale AI SEAL

Instruction following

Benchmarkno reasoninglowmediumhighxhighthinkingdefaultSourceTrend
LiveBench Languagenot in index82.6%#18LiveBench
IFBench48.4%#18065.9%#10074.0%#37Artificial Analysis

Human preference

Benchmarkno reasoninglowmediumhighxhighthinkingdefaultSourceTrend
LMArena Text1477#251466#46LMArena
EQ-Bench 41272#7EQ-Bench

Multimodal

Benchmarkno reasoninglowmediumhighxhighthinkingdefaultSourceTrend
LMArena Vision1298#141293#20LMArena
MMMU-Pro70.6%#11178.0%#5678.4%#54Artificial Analysis
VISTA50.9%#8Scale AI SEAL

Long context

Benchmarkno reasoninglowmediumhighxhighthinkingdefaultSourceTrend
AA-LCR58.3%#22576.7%#9682.0%#25Artificial Analysis

Composite

Benchmarkno reasoninglowmediumhighxhighthinkingdefaultSourceTrend
Epoch Capabilities Indexnot in index156.9#13Epoch AI Benchmarking Hub
LiveBench78.0%#12LiveBench
AA Intelligence Index18.2#19227.6#9739.0#42Artificial Analysis

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
OpenAI53 tok/s2.13 s$2.50$15.001.1M
Azure51 tok/s4.28 s$2.50$15.001.1M
Amazon Bedrock (US)32 tok/s1.33 s$2.75$16.501.1M
OpenAI Flex2 tok/s12 s$1.25$7.501.1M

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$4.54$9.08$13.61$18.15Jul 26Aug 26
input outputnow $2.75 in · $16.50 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.250 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0055$0.00482.1 min
Summarise a 30-page report12,000 / 600$0.039$0.0192.2 min
Code edit6,000 / 1,500$0.037$0.0272.3 min
Agentic coding session60,000 / 4,000$0.210$0.1092.6 min
Structured extraction2,000 / 200$0.0080$0.00462.1 min

See also

Data as of 9 Sept 2026. Compare these configurations.