BenchLeader
OpenAIReasoning model

GPT-5.2

Best configuration ranks #65 of 610 on the BenchLeader Index at 61.5 ±3.6. Last measured 8 Sept 2026. Released 11 Dec 2025.

Blended price
$4.81/M
$1.75 in · $14.00 out
Output speed
72 tok/s
First answer
138 s
first token 2.31 s
Context
400k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index62
  2. Reasoning61
  3. Coding53
  4. Agents & tools61
  5. Knowledge61
  6. Instruction following67
  7. Human preference68
  8. Multimodal60
  9. Long context68
  10. Composite67

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning50.1#29265 tok/s0.97 s$0.0049Agents & tools 40 · Coding 50 · Composite 50 · Instruction following 49 · Knowledge 58 · Long context 49 · Maths 51 · Multimodal 50 · Reasoning 54
low50.2#29072 tok/s138 s$0.0049Coding 50 · Knowledge 44 · Maths 59 · Reasoning 49
medium58.5#12372 tok/s138 s$0.0049Agents & tools 62 · Coding 59 · Composite 62 · Instruction following 64 · Knowledge 54 · Long context 62 · Maths 65 · Multimodal 59 · Reasoning 55
high55.5#17572 tok/s138 s$0.0049Agents & tools 55 · Coding 61 · Composite 50 · Human preference 63 · Knowledge 45 · Maths 62 · Multimodal 59 · Reasoning 56
xhigh57.3#14172 tok/s138 s$0.0049Agents & tools 51 · Coding 65 · Knowledge 56 · Maths 58 · Reasoning 62
defaultbest61.5#6572 tok/s138 s$0.0049Agents & tools 61 · Coding 53 · Composite 67 · Human preference 68 · Instruction following 67 · Knowledge 61 · Long context 68 · Multimodal 60 · Reasoning 61

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
GPQA Diamond73.2%#14582.7%#9987.9%#5388.2%#5291.4%#26Epoch AI Benchmarking Hub
Humanity's Last Exam27.8%#10Scale AI / CAIS
SimpleBench45.8%#5245.8%#52SimpleBench
LMArena Hard Prompts1460#801497#30LMArena
LiveBench Reasoningnot in index83.2%#33LiveBench
GPQA Diamond (AA)not in index71.2%#27886.4%#11190.3%#63Artificial Analysis
Humanity's Last Exam (AA)not in index8.0%#28226.7%#12937.7%#67Artificial Analysis
GPQA Diamond (Vals)not in index91.7%#18Vals AI
Kagi LLM Benchmark73.3%#18Kagi LLM Benchmark
ARC-AGI-155.7%#11772.7%#9678.7%#8486.2%#6894.5%#24ARC Prize
ARC-AGI-29.7%#10726.7%#9243.3%#7752.9%#6972.9%#33ARC Prize

Coding

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
SWE-bench Verified (Epoch)73.8%#17Epoch AI Benchmarking Hub
WeirdML49.6%#6749.6%#6763.4%#3972.2%#27WeirdML
GSO-Bench27.4%#9GSO-Bench
LMArena Coding1490#781515#35LMArena
LMArena WebDev1417#65LMArena
LiveBench Codingnot in index76.1%#33LiveBench
LiveCodeBench85.4%#31Vals AI
IOI54.8%#8Vals AI
SWE-bench (Vals)not in index75.8%#41Vals AI
SWE-Bench Pro29.9%#16Scale AI SEAL
SWE-bench Verified (bash only)72.8%#769.0%#14SWE-bench
SWE-bench Verified (any scaffold)not in index72.8%#11SWE-bench

Agents & tools

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
Terminal-Bench64.9%#964.9%#9Terminal-Bench
GDPval49.7%#1OpenAI
Remote Labor Index2.5%#72.1%#8Scale AI / CAIS
APEX-Agents23.0%#3334.4%#1923.0%#33Mercor
LiveBench Agentic Codingnot in index50.3%#31LiveBench
Terminal-Bench Hard31.8%#9943.2%#3947.0%#27Artificial Analysis
τ²-Bench Telecom (AA)not in index46.5%#19974.3%#13684.8%#95Artificial Analysis
MCP Atlas67.6%#21Scale AI SEAL
BFCL Overall55.9%#14Berkeley Function Calling Leaderboard
τ²-bench61.6%#884.8%#3τ²-bench

Maths

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
FrontierMath Tiers 1–367.4%#20Epoch AI Benchmarking Hub
FrontierMath Tier 431.7%#27Epoch AI Benchmarking Hub
OTIS Mock AIME62.2%#14378.9%#10593.9%#4296.1%#2596.1%#27Epoch AI Benchmarking Hub
ProofBench15.0%#44Vals AI
LiveBench Mathematicsnot in index93.2%#13LiveBench
AIME (Vals)96.9%#2Vals AI
MGSM94.0%#6Vals AI
AIME 202698.3%#4MathArena
HMMT February 202697.0%#3MathArena
MathArena Apex13.5%#18MathArena

Knowledge

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
SimpleQA Verified32.8%#5132.7%#5334.3%#4337.1%#39Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index78.2%#19LiveBench
AA-Omniscience-12.3#1540.3#92-0.9#102Artificial Analysis
MMLU-Pro86.2%#44Vals AI
LegalBench82.8%#58Vals AI
CorpFin65.9%#25Vals AI
TaxEval75.8%#11Vals AI
MedQA94.1%#19Vals AI
MultiNRC42.2%#17Scale AI SEAL

Instruction following

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
LiveBench Languagenot in index79.8%#23LiveBench
IFBench47.4%#18565.2%#10275.4%#31Artificial Analysis
TutorBench53.5%#11Scale AI SEAL

Human preference

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
LMArena Text1438#831476#26LMArena

Multimodal

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
LMArena Vision1247#521268#35LMArena
MMMU-Pro65.8%#14074.6%#82Artificial Analysis
VISTA46.6%#20Scale AI SEAL

Long context

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
AA-LCR45.7%#26670.3%#15782.7%#20Artificial Analysis

Composite

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
Epoch Capabilities Indexnot in index153.5#32Epoch AI Benchmarking Hub
LiveBench74.6%#30LiveBench
AA Intelligence Index17.0#20226.5#10530.4#77Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
OpenAI Fast69 tok/s2.31 s$3.50$28.00400k
OpenAI43 tok/s1.77 s$1.75$14.00400k
Azure34 tok/s4.61 s$1.75$14.00400k

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.175 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0049$0.00442.4 min
Summarise a 30-page report12,000 / 600$0.029$0.0152.4 min
Code edit6,000 / 1,500$0.032$0.0242.6 min
Agentic coding session60,000 / 4,000$0.161$0.0903.2 min
Structured extraction2,000 / 200$0.0063$0.00392.3 min

See also

Data as of 9 Sept 2026. Compare these configurations.