BenchLeader

GPT-5.5 vs GPT-5.6 Terra

Verdict
  • GPT-5.5 leads on quality: 66.6 vs 64.2.
  • GPT-5.5 is stronger in composite, human preference, instruction following, knowledge, long context, multimodal, reasoning.
  • GPT-5.6 Terra (xhigh) is stronger in agents & tools, coding, maths.
  • GPT-5.6 Terra (xhigh) is 2.5× cheaper ($4.50 vs $11.25 per 1M blended).
MetricGPT-5.5GPT-5.6 Terra (xhigh)
BenchLeader Index66.664.2
Agents & tools score68.469.6
Coding score57.363.7
Composite score77.777.1
Human preference score67.666.3
Instruction following score73.365.0
Knowledge score73.760.9
Long context score68.866.0
Multimodal score64.963.0
Reasoning score69.758.2
Maths score71.0
Blended price $/M$11.25$4.50
Output speed88 tok/s93 tok/s
Time to first answer62.3 s32.5 s
Context window1.1M1M
Terminal-Bench84.7%
SimpleBench69.0%48.9%
SciCode51.6%
Remote Labor Index6.3%
APEX-Agents38.5%
FrontierCode43.0%
ProofBench74.0%
Epoch Capabilities Index159.1
LMArena Text14771466
LMArena Hard Prompts14981491
LMArena Coding15091519
LMArena WebDev14581521
LMArena Vision12961269
LMArena Agent31.5
AA Intelligence Index38.638.2
IFBench75.8%66.3%
AA-LCR84.3%79.0%
MMMU-Pro79.9%79.5%
AA-Omniscience20.5-3.0
Terminal-Bench Hard60.6%62.9%
GPQA Diamond (AA)93.5%90.8%
Humanity's Last Exam (AA)45.8%41.9%
SciCode (AA)55.8%52.3%
τ²-Bench Telecom (AA)93.9%80.4%
LiveCodeBench85.9%
MMLU-Pro86.7%
IOI65.3%
CorpFin65.3%
TaxEval76.2%
GPQA Diamond (Vals)90.9%
HiL-Bench39.7%
EQ-Bench 41315
Kagi LLM Benchmark88.8%
ARC-AGI-194.0%
ARC-AGI-274.2%
ARC-AGI-30.7%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.