BenchLeader
OpenAIReasoning modelAuto-detected

GPT-5.4 Pro

Best configuration ranks #36 of 610 on the BenchLeader Index at 64.2 ±4.0. Last measured 23 Mar 2026. Released 5 Mar 2026.

Blended price
$67.50/M
$30.00 in · $180.00 out
Output speed
1 tok/s
First answer
6.70 s
Context
1.1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index64
  2. Reasoning75
  3. Knowledge73
  4. Instruction following67
  5. Multimodal70

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning1 tok/s6.70 s$0.066Coding 55
xhigh61.1#731 tok/s6.70 s$0.066Knowledge 55 · Maths 73 · Reasoning 69
defaultbest64.2#361 tok/s6.70 s$0.066Instruction following 67 · Knowledge 73 · Multimodal 70 · Reasoning 75

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoningxhighdefaultSource
GPQA Diamond94.6%#4Epoch AI Benchmarking Hub
Humanity's Last Exam44.3%#3Scale AI / CAIS
SimpleBench74.1%#10SimpleBench
ARC-AGI-194.5%#24ARC Prize
ARC-AGI-283.3%#24ARC Prize

Coding

Benchmarkno reasoningxhighdefaultSource
WeirdML57.4%#50WeirdML

Maths

Benchmarkno reasoningxhighdefaultSource
FrontierMath Tiers 1–382.5%#9Epoch AI Benchmarking Hub
FrontierMath Tier 458.5%#17Epoch AI Benchmarking Hub
MathArena Apex69.8%#3MathArena

Knowledge

Benchmarkno reasoningxhighdefaultSource
SimpleQA Verified46.3%#30Epoch AI Benchmarking Hub
MultiNRC62.3%#4Scale AI SEAL

Instruction following

Benchmarkno reasoningxhighdefaultSource
MultiChallenge69.2%#4Scale AI SEAL
TutorBench56.6%#2Scale AI SEAL

Multimodal

Benchmarkno reasoningxhighdefaultSource
VISTA53.9%#2Scale AI SEAL

Composite

Benchmarkno reasoningxhighdefaultSource
Epoch Capabilities Indexnot in index158.9#9Epoch AI Benchmarking Hub

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
OpenAI1 tok/s6.70 s$30.00$180.001.1M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0665.1 min
Summarise a 30-page report12,000 / 600$0.46810.1 min
Code edit6,000 / 1,500$0.45025.1 min
Agentic coding session60,000 / 4,000$2.5266.8 min
Structured extraction2,000 / 200$0.0963.4 min

See also

Data as of 9 Sept 2026. Compare these configurations.