BenchLeader

GPT-5.4 Pro vs GPT-5.5

Verdict
  • GPT-5.5 leads on quality: 66.6 vs 64.2.
  • GPT-5.4 Pro is stronger in multimodal, reasoning.
  • GPT-5.5 is stronger in instruction following, knowledge, agents & tools, coding, composite, human preference, long context.
  • GPT-5.5 is 6.0× cheaper ($11.25 vs $67.50 per 1M blended).
  • GPT-5.5 streams 87.7× faster (88 vs 1 tokens per second).
MetricGPT-5.4 ProGPT-5.5
BenchLeader Index64.266.6
Instruction following score67.373.3
Knowledge score73.273.7
Multimodal score69.864.9
Reasoning score75.169.7
Agents & tools score68.4
Coding score57.3
Composite score77.7
Human preference score67.6
Long context score68.8
Blended price $/M$67.50$11.25
Output speed1 tok/s88 tok/s
Time to first answer6.7 s62.3 s
Context window1.1M1.1M
Humanity's Last Exam44.3%
Terminal-Bench84.7%
SimpleBench74.1%69.0%
Remote Labor Index6.3%
APEX-Agents38.5%
FrontierCode43.0%
Epoch Capabilities Index158.9159.1
LMArena Text1477
LMArena Hard Prompts1498
LMArena Coding1509
LMArena WebDev1458
LMArena Vision1296
LMArena Agent3
AA Intelligence Index38.6
IFBench75.8%
AA-LCR84.3%
MMMU-Pro79.9%
AA-Omniscience20.5
Terminal-Bench Hard60.6%
GPQA Diamond (AA)93.5%
Humanity's Last Exam (AA)45.8%
SciCode (AA)55.8%
τ²-Bench Telecom (AA)93.9%
MultiChallenge69.2%
VISTA53.9%
MultiNRC62.3%
HiL-Bench39.7%
TutorBench56.6%
EQ-Bench 41315
Kagi LLM Benchmark88.8%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.