BenchLeader

GPT-5 Pro vs o3

Verdict
  • GPT-5 Pro and o3 are level on quality (60.8 vs 61.6).
  • GPT-5 Pro is stronger in knowledge, multimodal.
  • o3 is stronger in reasoning, agents & tools, coding, composite, human preference, instruction following, long context, maths.
  • o3 is 12× cheaper ($3.50 vs $41.25 per 1M blended).
MetricGPT-5 Proo3
BenchLeader Index60.861.6
Knowledge score68.058.0
Multimodal score67.457.9
Reasoning score57.661.5
Agents & tools score70.3
Coding score62.0
Composite score54.5
Human preference score62.3
Instruction following score65.0
Long context score63.8
Maths score78.6
Blended price $/M$41.25$3.50
Output speed107 tok/s
Time to first answer6.3 s
Context window400k200k
Humanity's Last Exam31.6%
SimpleBench61.6%
Epoch Capabilities Index150.3146.9
LMArena Text1432
LMArena Hard Prompts1441
LMArena Coding1460
LMArena Vision1214
AA Intelligence Index20.2
IFBench71.4%
AA-LCR74.7%
MMMU-Pro70.1%
AA-Omniscience-15.6
Terminal-Bench Hard37.1%
GPQA Diamond (AA)82.7%
Humanity's Last Exam (AA)20.1%
τ²-Bench Telecom (AA)80.7%
PRBench Finance51.1%47.7%
PRBench Legal49.9%48.6%
VISTA52.4%
MultiNRC65.2%
MMMU (validation)82.9%
MMMU-Pro (official)76.4%
Kagi LLM Benchmark76.8%67.6%
IFEval (HELM)86.9%
Omni-MATH (HELM)71.4%
WildBench (HELM)86.1%
MMLU-Pro (HELM)85.9%
GPQA Diamond (HELM)75.3%
HELM Capabilities mean81.1%
Aider Polyglot81.3%
SWE-bench Verified (bash only)58.4%
SWE-bench Verified (any scaffold)58.4%
ARC-AGI-170.2%
ARC-AGI-218.3%
BFCL Overall63.0%

Data as of 2026-09-12. Best configuration of each model; every score links to its source on the model pages.

GPT-5 Pro vs o3: questions

Is GPT-5 Pro better than o3?
GPT-5 Pro and o3 are level on quality (60.8 vs 61.6). The BenchLeader Index combines every independent quality benchmark; o3 is ahead overall as of 2026-09-12, but check the category scores for your use.
Which is cheaper, GPT-5 Pro or o3?
o3 is cheaper: $3.50 against $41.25 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
GPT-5 Pro accepts more context: 400k against 200k tokens.