BenchLeader
OpenAIReasoning modelAuto-detected

GPT-5.2 Codex

Best configuration ranks #77 of 610 on the BenchLeader Index at 60.7 ±4.9. Last measured 8 Sept 2026. Released 18 Dec 2025.

Blended price
$4.81/M
$1.75 in · $14.00 out
Output speed
25 tok/s
First answer
6.74 s
Context
400k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index61
  2. Coding56
  3. Agents & tools64
  4. Knowledge63
  5. Instruction following75
  6. Multimodal61
  7. Long context68
  8. Composite57

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
high25 tok/s6.74 s$0.0049
defaultbest60.7#7725 tok/s6.74 s$0.0049Agents & tools 64 · Coding 56 · Composite 57 · Instruction following 75 · Knowledge 63 · Long context 68 · Multimodal 61

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkhighdefaultSourceTrend
LiveBench Reasoningnot in index77.7%#42LiveBench
GPQA Diamond (AA)not in index89.9%#67Artificial Analysis
Humanity's Last Exam (AA)not in index35.7%#75Artificial Analysis

Coding

BenchmarkhighdefaultSourceTrend
LMArena WebDev1339#94LMArena
LiveBench Codingnot in index83.6%#4LiveBench
LiveCodeBench88.0%#9Vals AI
SWE-bench (Vals)not in index72.4%#53Vals AI
SWE-Bench Pro41.0%#11Scale AI SEAL
SWE-bench Verified (bash only)72.8%#7SWE-bench
SWE-bench Verified (any scaffold)not in index72.8%#11SWE-bench

Agents & tools

BenchmarkhighdefaultSourceTrend
Terminal-Bench66.5%#8Terminal-Bench
APEX-Agents27.6%#28Mercor
LiveBench Agentic Codingnot in index49.4%#32LiveBench
Terminal-Bench Hard37.1%#67Artificial Analysis
τ²-Bench Telecom (AA)not in index92.1%#53Artificial Analysis

Maths

BenchmarkhighdefaultSourceTrend
LiveBench Mathematicsnot in index88.8%#28LiveBench

Knowledge

BenchmarkhighdefaultSourceTrend
LiveBench Data Analysisnot in index78.2%#18LiveBench
AA-Omniscience-2.2#105Artificial Analysis

Instruction following

BenchmarkhighdefaultSourceTrend
LiveBench Languagenot in index73.7%#43LiveBench
IFBench77.6%#16Artificial Analysis

Multimodal

BenchmarkhighdefaultSourceTrend
MMMU-Pro76.3%#64Artificial Analysis

Long context

BenchmarkhighdefaultSourceTrend
AA-LCR82.3%#21Artificial Analysis

Composite

BenchmarkhighdefaultSourceTrend
LiveBench74.0%#33LiveBench
AA Intelligence Index28.5#88Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Azure25 tok/s6.74 s$1.75$14.00400k

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.175 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0049$0.004418.7 s
Summarise a 30-page report12,000 / 600$0.029$0.01530.7 s
Code edit6,000 / 1,500$0.032$0.0241.1 min
Agentic coding session60,000 / 4,000$0.161$0.0902.8 min
Structured extraction2,000 / 200$0.0063$0.003914.7 s

See also

Data as of 9 Sept 2026. Compare these configurations.