BenchLeader

GPT-5.2-Codex vs o3

Verdict
  • GPT-5.2-Codex and o3 are level on quality (60.8 vs 61.6).
  • GPT-5.2-Codex is stronger in composite, instruction following, knowledge, long context, multimodal.
  • o3 is stronger in agents & tools, coding, human preference, maths, reasoning.
  • o3 is 1.4× cheaper ($3.50 vs $4.81 per 1M blended).
  • o3 streams 4.9× faster (107 vs 22 tokens per second).
MetricGPT-5.2-Codexo3
BenchLeader Index60.861.6
Agents & tools score63.570.3
Coding score55.762.0
Composite score56.854.5
Instruction following score75.165.0
Knowledge score63.258.0
Long context score67.763.8
Multimodal score60.857.9
Human preference score62.3
Maths score78.6
Reasoning score61.5
Blended price $/M$4.81$3.50
Output speed22 tok/s107 tok/s
Time to first answer5.9 s6.3 s
Context window400k200k
Terminal-Bench66.5%
APEX-Agents27.6%
Epoch Capabilities Index146.9
LMArena Text1432
LMArena Hard Prompts1441
LMArena Coding1460
LMArena WebDev1339
LMArena Vision1214
LiveBench74.0%
LiveBench Reasoning77.7%
LiveBench Coding83.6%
LiveBench Agentic Coding49.4%
LiveBench Mathematics88.8%
LiveBench Data Analysis78.2%
LiveBench Language73.7%
AA Intelligence Index28.520.2
IFBench77.6%71.4%
AA-LCR82.3%74.7%
MMMU-Pro76.3%70.1%
AA-Omniscience-2.2-15.6
Terminal-Bench Hard37.1%37.1%
GPQA Diamond (AA)89.9%82.7%
Humanity's Last Exam (AA)35.7%20.1%
τ²-Bench Telecom (AA)92.1%80.7%
LiveCodeBench88.0%
SWE-Bench Pro41.0%
PRBench Finance47.7%
PRBench Legal48.6%
MMMU (validation)82.9%
MMMU-Pro (official)76.4%
Kagi LLM Benchmark67.6%
IFEval (HELM)86.9%
Omni-MATH (HELM)71.4%
WildBench (HELM)86.1%
MMLU-Pro (HELM)85.9%
GPQA Diamond (HELM)75.3%
HELM Capabilities mean81.1%
Aider Polyglot81.3%
SWE-bench Verified (bash only)72.8%58.4%
SWE-bench Verified (any scaffold)72.8%58.4%
BFCL Overall63.0%

Data as of 2026-09-12. Best configuration of each model; every score links to its source on the model pages.

GPT-5.2-Codex vs o3: questions

Is GPT-5.2-Codex better than o3?
GPT-5.2-Codex and o3 are level on quality (60.8 vs 61.6). The BenchLeader Index combines every independent quality benchmark; o3 is ahead overall as of 2026-09-12, but check the category scores for your use.
Is GPT-5.2-Codex better than o3 for coding?
o3 scores higher in coding (62 vs 56 on the category index, where 50 is average).
Is GPT-5.2-Codex better than o3 for agentic tasks?
o3 scores higher in agentic tasks (70 vs 64 on the category index, where 50 is average).
Which is cheaper, GPT-5.2-Codex or o3?
o3 is cheaper: $3.50 against $4.81 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.2-Codex or o3?
o3 streams faster: 107 against 22 output tokens per second.
Which has the larger context window?
GPT-5.2-Codex accepts more context: 400k against 200k tokens.