BenchLeader

GPT-5 vs o3

Verdict
  • GPT-5 and o3 are level on quality (60.8 vs 61.6).
  • GPT-5 is stronger in composite, knowledge, long context, multimodal, reasoning.
  • o3 is stronger in agents & tools, coding, human preference, maths.
  • They cost about the same ($3.44 per 1M blended).
MetricGPT-5o3
BenchLeader Index60.861.6
Agents & tools score51.070.3
Coding score60.962.0
Composite score58.154.5
Human preference score61.762.3
Instruction following score65.065.0
Knowledge score62.258.0
Long context score65.663.8
Maths score72.878.6
Multimodal score59.657.9
Reasoning score64.261.5
Blended price $/M$3.44$3.50
Output speed85 tok/s107 tok/s
Time to first answer60.7 s6.3 s
Context window400k200k
Terminal-Bench49.6%
SciCode42.9%
Remote Labor Index1.7%
WeirdML39.8%
APEX-Agents18.3%
Epoch Capabilities Index150146.9
LMArena Text14271432
LMArena Hard Prompts14491441
LMArena Coding14621460
LMArena Vision12321214
AA Intelligence Index23.020.2
IFBench73.1%71.4%
AA-LCR78.2%74.7%
MMMU-Pro74.2%70.1%
AA-Omniscience-8.7-15.6
Terminal-Bench Hard32.6%37.1%
GPQA Diamond (AA)85.3%82.7%
Humanity's Last Exam (AA)28.5%20.1%
τ²-Bench Telecom (AA)84.8%80.7%
PRBench Finance51.3%47.7%
PRBench Legal49.0%48.6%
VISTA49.7%
MultiNRC52.1%
TutorBench55.3%
MMMU (validation)82.9%
MMMU-Pro (official)76.4%
Kagi LLM Benchmark72.7%67.6%
IFEval (HELM)87.5%86.9%
Omni-MATH (HELM)64.7%71.4%
WildBench (HELM)85.7%86.1%
MMLU-Pro (HELM)86.3%85.9%
GPQA Diamond (HELM)79.1%75.3%
HELM Capabilities mean80.7%81.1%
Aider Polyglot88.0%81.3%
SWE-bench Verified (bash only)58.4%
SWE-bench Verified (any scaffold)75.6%58.4%
BFCL Overall63.0%

Data as of 2026-09-12. Best configuration of each model; every score links to its source on the model pages.

GPT-5 vs o3: questions

Is GPT-5 better than o3?
GPT-5 and o3 are level on quality (60.8 vs 61.6). The BenchLeader Index combines every independent quality benchmark; o3 is ahead overall as of 2026-09-12, but check the category scores for your use.
Is GPT-5 better than o3 for coding?
o3 scores higher in coding (62 vs 61 on the category index, where 50 is average).
Is GPT-5 better than o3 for agentic tasks?
o3 scores higher in agentic tasks (70 vs 51 on the category index, where 50 is average).
Which is cheaper, GPT-5 or o3?
GPT-5 is cheaper: $3.44 against $3.50 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5 or o3?
o3 streams faster: 107 against 85 output tokens per second.
Which has the larger context window?
GPT-5 accepts more context: 400k against 200k tokens.