BenchLeader

GPT-5.4 Pro vs Grok 4.6

Verdict
  • GPT-5.4 Pro and Grok 4.6 (medium) are level on quality (64.2 vs 63.8).
  • GPT-5.4 Pro is stronger in instruction following, multimodal, reasoning.
  • Grok 4.6 (medium) is stronger in knowledge, coding, composite, long context.
  • Grok 4.6 (medium) is 23× cheaper ($3.00 vs $67.50 per 1M blended).
  • Grok 4.6 (medium) streams 54.8× faster (55 vs 1 tokens per second).
MetricGPT-5.4 ProGrok 4.6 (medium)
BenchLeader Index64.263.8
Instruction following score67.3
Knowledge score73.277.3
Multimodal score69.8
Reasoning score75.162.1
Coding score63.9
Composite score83.2
Long context score67.1
Blended price $/M$67.50$3.00
Output speed1 tok/s55 tok/s
Time to first answer6.7 s36.7 s
Context window1.1M500k
Humanity's Last Exam44.3%
SimpleBench74.1%
SciCode54.6%
Epoch Capabilities Index158.9
AA Intelligence Index43.0
AA-LCR81.0%
AA-Omniscience28
GPQA Diamond (AA)93.5%
Humanity's Last Exam (AA)42.1%
SciCode (AA)55.9%
MultiChallenge69.2%
VISTA53.9%
MultiNRC62.3%
TutorBench56.6%
ARC-AGI-187.5%
ARC-AGI-261.3%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.