BenchLeader

GPT-5.4 Pro vs GPT-5.6 Terra

Verdict
  • GPT-5.4 Pro and GPT-5.6 Terra (xhigh) are level on quality (64.2 vs 64.2).
  • GPT-5.4 Pro is stronger in instruction following, knowledge, multimodal, reasoning.
  • GPT-5.6 Terra (xhigh) is stronger in agents & tools, coding, composite, human preference, long context, maths.
  • GPT-5.6 Terra (xhigh) is 15× cheaper ($4.50 vs $67.50 per 1M blended).
  • GPT-5.6 Terra (xhigh) streams 93.1× faster (93 vs 1 tokens per second).
MetricGPT-5.4 ProGPT-5.6 Terra (xhigh)
BenchLeader Index64.264.2
Instruction following score67.365.0
Knowledge score73.260.9
Multimodal score69.863.0
Reasoning score75.158.2
Agents & tools score69.6
Coding score63.7
Composite score77.1
Human preference score66.3
Long context score66.0
Maths score71.0
Blended price $/M$67.50$4.50
Output speed1 tok/s93 tok/s
Time to first answer6.7 s32.5 s
Context window1.1M1M
Humanity's Last Exam44.3%
SimpleBench74.1%48.9%
SciCode51.6%
ProofBench74.0%
Epoch Capabilities Index158.9
LMArena Text1466
LMArena Hard Prompts1491
LMArena Coding1519
LMArena WebDev1521
LMArena Vision1269
LMArena Agent1.5
AA Intelligence Index38.2
IFBench66.3%
AA-LCR79.0%
MMMU-Pro79.5%
AA-Omniscience-3.0
Terminal-Bench Hard62.9%
GPQA Diamond (AA)90.8%
Humanity's Last Exam (AA)41.9%
SciCode (AA)52.3%
τ²-Bench Telecom (AA)80.4%
LiveCodeBench85.9%
MMLU-Pro86.7%
IOI65.3%
CorpFin65.3%
TaxEval76.2%
GPQA Diamond (Vals)90.9%
MultiChallenge69.2%
VISTA53.9%
MultiNRC62.3%
TutorBench56.6%
ARC-AGI-194.0%
ARC-AGI-274.2%
ARC-AGI-30.7%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.