BenchLeader

GPT-5.5 vs Grok 4.5

Verdict
  • GPT-5.5 leads on quality: 66.6 vs 63.4.
  • GPT-5.5 is stronger in agents & tools, composite, human preference, instruction following, long context, multimodal, reasoning.
  • Grok 4.5 is stronger in coding, knowledge.
  • Grok 4.5 is 3.8× cheaper ($3.00 vs $11.25 per 1M blended).
  • GPT-5.5 streams 1.6× faster (88 vs 56 tokens per second).
MetricGPT-5.5Grok 4.5
BenchLeader Index66.663.4
Agents & tools score68.458.6
Coding score57.359.9
Composite score77.766.0
Human preference score67.666.9
Instruction following score73.3
Knowledge score73.776.0
Long context score68.866.2
Multimodal score64.964.8
Reasoning score69.768.9
Blended price $/M$11.25$3.00
Output speed88 tok/s56 tok/s
Time to first answer62.3 s12.7 s
Context window1.1M500k
Terminal-Bench84.7%
SimpleBench69.0%70.0%
Remote Labor Index6.3%
WeirdML46.4%
APEX-Agents38.5%34.2%
FrontierCode43.0%42.4%
Epoch Capabilities Index159.1153.9
LMArena Text14771471
LMArena Hard Prompts14981495
LMArena Coding15091523
LMArena WebDev14581556
LMArena Vision12961291
LMArena Agent33.9
LiveBench75.8%
LiveBench Reasoning87.2%
LiveBench Coding68.6%
LiveBench Agentic Coding56.5%
LiveBench Mathematics90.8%
LiveBench Data Analysis73.0%
LiveBench Language82.8%
AA Intelligence Index38.639.1
IFBench75.8%
AA-LCR84.3%79.3%
MMMU-Pro79.9%80.4%
AA-Omniscience20.525.3
Terminal-Bench Hard60.6%
GPQA Diamond (AA)93.5%93.1%
Humanity's Last Exam (AA)45.8%42.7%
SciCode (AA)55.8%55.0%
τ²-Bench Telecom (AA)93.9%
HiL-Bench39.7%
EQ-Bench 41315
Kagi LLM Benchmark88.8%83.5%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.