BenchLeader

GPT-5.5 Instant vs Grok 4.5

Verdict
  • Grok 4.5 leads on quality: 63.5 vs 60.5.
  • GPT-5.5 Instant is stronger in agents & tools, coding, human preference, instruction following, maths.
  • Grok 4.5 is stronger in composite, knowledge, long context, multimodal, reasoning.
  • Grok 4.5 is 3.8× cheaper ($3.00 vs $11.25 per 1M blended).
  • GPT-5.5 Instant streams 2.5× faster (129 vs 52 tokens per second).
MetricGPT-5.5 InstantGrok 4.5
BenchLeader Index60.563.5
Agents & tools score70.758.6
Coding score61.159.8
Composite score62.966.2
Human preference score67.567.2
Instruction following score69.8
Knowledge score69.476.2
Long context score61.466.2
Maths score42.5
Multimodal score59.764.9
Reasoning score62.669.2
Blended price $/M$11.25$3.00
Output speed129 tok/s52 tok/s
Time to first answer16.8 s13.6 s
Context window400k500k
GPQA Diamond82.5%
FrontierMath Tiers 1–326.3%
FrontierMath Tier 42.4%
OTIS Mock AIME68.1%
SimpleBench70.0%
SciCode48.6%
WeirdML46.4%
APEX-Agents34.2%
FrontierCode42.4%
Epoch Capabilities Index142.5153.9
LMArena Text14741471
LMArena Hard Prompts14911495
LMArena Coding15141523
LMArena WebDev1556
LMArena Vision12511291
LMArena Agent3.9
LiveBench75.8%
LiveBench Reasoning87.2%
LiveBench Coding68.6%
LiveBench Agentic Coding56.5%
LiveBench Mathematics90.8%
LiveBench Data Analysis73.0%
LiveBench Language82.8%
AA Intelligence Index26.839.1
IFBench71.5%
AA-LCR70.0%79.3%
MMMU-Pro80.4%
AA-Omniscience11.225.3
Terminal-Bench Hard42.4%
GPQA Diamond (AA)84.7%93.1%
Humanity's Last Exam (AA)21.6%42.7%
SciCode (AA)52.5%55.0%
τ²-Bench Telecom (AA)49.4%
Kagi LLM Benchmark83.5%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

GPT-5.5 Instant vs Grok 4.5: questions

Is GPT-5.5 Instant better than Grok 4.5?
Grok 4.5 leads on quality: 63.5 vs 60.5. The BenchLeader Index combines every independent quality benchmark; Grok 4.5 is ahead overall as of 2026-09-10, but check the category scores for your use.
Is GPT-5.5 Instant better than Grok 4.5 for coding?
GPT-5.5 Instant scores higher in coding (61 vs 60 on the category index, where 50 is average).
Is GPT-5.5 Instant better than Grok 4.5 for agentic tasks?
GPT-5.5 Instant scores higher in agentic tasks (71 vs 59 on the category index, where 50 is average).
Which is cheaper, GPT-5.5 Instant or Grok 4.5?
Grok 4.5 is cheaper: $3.00 against $11.25 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.5 Instant or Grok 4.5?
GPT-5.5 Instant streams faster: 129 against 52 output tokens per second.
Which has the larger context window?
Grok 4.5 accepts more context: 500k against 400k tokens.