BenchLeader

GPT-5.4 vs o3

Verdict
  • GPT-5.4 leads on quality: 63.5 vs 61.6.
  • GPT-5.4 is stronger in composite, human preference, instruction following, knowledge, long context, multimodal, reasoning.
  • o3 is stronger in agents & tools, coding, maths.
  • o3 is 1.6× cheaper ($3.50 vs $5.63 per 1M blended).
MetricGPT-5.4o3
BenchLeader Index63.561.6
Agents & tools score66.270.3
Coding score51.662.0
Composite score78.454.5
Human preference score64.762.3
Instruction following score71.965.0
Knowledge score67.058.0
Long context score67.663.8
Multimodal score64.057.9
Reasoning score62.661.5
Maths score78.6
Blended price $/M$5.63$3.50
Output speed132 tok/s107 tok/s
Time to first answer102.0 s6.3 s
Context window1.1M200k
Terminal-Bench81.8%
APEX-Agents34.9%
Epoch Capabilities Index156.9146.9
LMArena Text14661432
LMArena Hard Prompts14881441
LMArena Coding15131460
LMArena WebDev1387
LMArena Vision12931214
AA Intelligence Index39.020.2
IFBench74.0%71.4%
AA-LCR82.0%74.7%
MMMU-Pro78.4%70.1%
AA-Omniscience5.8-15.6
Terminal-Bench Hard57.6%37.1%
GPQA Diamond (AA)92.0%82.7%
Humanity's Last Exam (AA)43.7%20.1%
τ²-Bench Telecom (AA)87.1%80.7%
PRBench Finance47.7%
PRBench Legal48.6%
HiL-Bench9.7%
EQ-Bench 41272
MMMU (validation)82.9%
MMMU-Pro (official)76.4%
Kagi LLM Benchmark63.8%67.6%
IFEval (HELM)86.9%
Omni-MATH (HELM)71.4%
WildBench (HELM)86.1%
MMLU-Pro (HELM)85.9%
GPQA Diamond (HELM)75.3%
HELM Capabilities mean81.1%
Aider Polyglot81.3%
SWE-bench Verified (bash only)58.4%
SWE-bench Verified (any scaffold)58.4%
BFCL Overall63.0%

Data as of 2026-09-12. Best configuration of each model; every score links to its source on the model pages.

GPT-5.4 vs o3: questions

Is GPT-5.4 better than o3?
GPT-5.4 leads on quality: 63.5 vs 61.6. The BenchLeader Index combines every independent quality benchmark; GPT-5.4 is ahead overall as of 2026-09-12, but check the category scores for your use.
Is GPT-5.4 better than o3 for coding?
o3 scores higher in coding (62 vs 52 on the category index, where 50 is average).
Is GPT-5.4 better than o3 for agentic tasks?
o3 scores higher in agentic tasks (70 vs 66 on the category index, where 50 is average).
Which is cheaper, GPT-5.4 or o3?
o3 is cheaper: $3.50 against $5.63 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.4 or o3?
GPT-5.4 streams faster: 132 against 107 output tokens per second.
Which has the larger context window?
GPT-5.4 accepts more context: 1.1M against 200k tokens.