BenchLeader
OpenAIReasoning model

GPT-5.1-Codex-Max

Not yet ranked: too few independent results so far. Released 13 Nov 2025.

Blended price
$3.44/M
$1.25 in · $10.00 out
Output speed
29 tok/s
First answer
4.29 s
Context
400k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Coding57
  2. Agents & tools63
  3. Maths36

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Coding

BenchmarkmaxSource
LiveCodeBench83.6%#46Vals AI
IOI21.4%#22Vals AI

Agents & tools

BenchmarkmaxSource
Terminal-Bench60.4%#14Terminal-Bench

Maths

BenchmarkmaxSource
ProofBench9.0%#50Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Azure29 tok/s4.29 s$1.25$10.00400k

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.125 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0035$0.003214.8 s
Summarise a 30-page report12,000 / 600$0.021$0.01125.3 s
Code edit6,000 / 1,500$0.022$0.01756.9 s
Agentic coding session60,000 / 4,000$0.115$0.0642.4 min
Structured extraction2,000 / 200$0.0045$0.002811.3 s

See also

Data as of 9 Sept 2026. Compare with another model.