BenchLeader

GPT-5.4 Pro vs GPT-5.6 Sol

Verdict
  • GPT-5.6 Sol (high) leads on quality: 68.0 vs 64.2.
  • GPT-5.4 Pro is stronger in multimodal, reasoning.
  • GPT-5.6 Sol (high) is stronger in instruction following, knowledge, agents & tools, coding, composite, long context.
  • GPT-5.6 Sol (high) is 8.4× cheaper ($8.00 vs $67.50 per 1M blended).
  • GPT-5.6 Sol (high) streams 59.0× faster (59 vs 1 tokens per second).
MetricGPT-5.4 ProGPT-5.6 Sol (high)
BenchLeader Index64.268.0
Instruction following score67.367.6
Knowledge score73.273.7
Multimodal score69.866.5
Reasoning score75.162.9
Agents & tools score87.6
Coding score72.0
Composite score82.5
Long context score67.4
Blended price $/M$67.50$8.00
Output speed1 tok/s59 tok/s
Time to first answer6.7 s31.4 s
Context window1.1M1M
Humanity's Last Exam44.3%
SimpleBench74.1%
SciCode56.9%
WeirdML88.8%
Epoch Capabilities Index158.9
AA Intelligence Index42.5
IFBench69.2%
AA-LCR81.7%
MMMU-Pro81.8%
AA-Omniscience20.4
Terminal-Bench Hard62.1%
GPQA Diamond (AA)92.8%
Humanity's Last Exam (AA)46.0%
SciCode (AA)57.8%
τ²-Bench Telecom (AA)83.3%
MultiChallenge69.2%
EnigmaEval37.1%
VISTA53.9%
MultiNRC62.3%
TutorBench56.6%
ARC-AGI-197.0%
ARC-AGI-285.4%
ARC-AGI-32.2%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.