BenchLeader

GPT-5.2 vs Step 5 Preview

Verdict
  • Step 5 Preview leads on quality: 65.0 vs 62.0.
  • GPT-5.2 (xhigh) is stronger in agents & tools, coding, instruction following, maths.
  • Step 5 Preview is stronger in composite, knowledge, long context, reasoning, multimodal.
  • Step 5 Preview is 3.4× cheaper ($1.43 vs $4.81 per 1M blended).
MetricGPT-5.2 (xhigh)Step 5 Preview
BenchLeader Index62.065.0
Agents & tools score56.4
Coding score63.2
Composite score67.684.4
Instruction following score73.4
Knowledge score57.972.0
Long context score67.870.8
Maths score58.4
Reasoning score62.882.1
Multimodal score60.8
Blended price $/M$4.81$1.43
Output speed67 tok/s
Time to first answer94.0 s
Context window400k1M
GPQA Diamond91.4%
FrontierMath Tiers 1–367.4%
FrontierMath Tier 431.7%
OTIS Mock AIME96.1%
SimpleQA Verified37.1%
WeirdML72.2%
APEX-Agents34.4%
ProofBench15.0%
AA Intelligence Index30.443.6
IFBench75.4%
AA-LCR82.7%88.3%
MMMU-Pro76.4%
AA-Omniscience-0.916.4
Terminal-Bench Hard47.0%
GPQA Diamond (AA)90.3%
Humanity's Last Exam (AA)37.7%46.5%
SciCode (AA)58.9%
τ²-Bench Telecom (AA)84.8%
AIME (Vals)96.9%
LiveCodeBench85.4%
MMLU-Pro86.2%
LegalBench82.8%
CorpFin65.9%
TaxEval75.8%
MedQA94.1%
MGSM94.0%
SWE-bench (Vals)75.8%
GPQA Diamond (Vals)91.7%
MCP Atlas67.6%
ARC-AGI-186.2%
ARC-AGI-252.9%
CritPt11.6%20.9%
GDPval (AA)53.6%
CaseLaw v266.0%
MedCode49.8%
MedScribe84.4%
MMMU-Pro (Vals)86.7%
MortgageTax67.1%
SAGE49.3%
Vibe Code Bench v1.153.5%
DrugDiscoveryBench29.3%
Chess Puzzles49.0%
EBR-bench23.0%
VPCT84.0%
LMCA43.9%
DTBench90.9%

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

GPT-5.2 vs Step 5 Preview: questions

Is GPT-5.2 better than Step 5 Preview?
Step 5 Preview leads on quality: 65.0 vs 62.0. The BenchLeader Index combines every independent quality benchmark; Step 5 Preview is ahead overall as of 2026-09-19, but check the category scores for your use.
Which is cheaper, GPT-5.2 or Step 5 Preview?
Step 5 Preview is cheaper: $1.43 against $4.81 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Step 5 Preview accepts more context: 1M against 400k tokens.