BenchLeader

GPT-5 Pro vs Grok 4.6

Verdict
  • Grok 4.6 (medium) leads on quality: 63.9 vs 60.8.
  • GPT-5 Pro is stronger in multimodal.
  • Grok 4.6 (medium) is stronger in knowledge, reasoning, coding, composite, long context.
  • Grok 4.6 (medium) is 14× cheaper ($3.00 vs $41.25 per 1M blended).
MetricGPT-5 ProGrok 4.6 (medium)
BenchLeader Index60.863.9
Knowledge score68.077.5
Multimodal score67.4
Reasoning score57.662.1
Coding score64.1
Composite score83.4
Long context score67.1
Blended price $/M$41.25$3.00
Output speed53 tok/s
Time to first answer33.4 s
Context window400k500k
Humanity's Last Exam31.6%
SimpleBench61.6%
SciCode54.6%
Epoch Capabilities Index150.3
AA Intelligence Index43.0
AA-LCR81.0%
AA-Omniscience28
GPQA Diamond (AA)93.5%
Humanity's Last Exam (AA)42.1%
SciCode (AA)55.9%
PRBench Finance51.1%
PRBench Legal49.9%
VISTA52.4%
MultiNRC65.2%
Kagi LLM Benchmark76.8%
ARC-AGI-170.2%87.5%
ARC-AGI-218.3%61.3%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

GPT-5 Pro vs Grok 4.6: questions

Is GPT-5 Pro better than Grok 4.6?
Grok 4.6 (medium) leads on quality: 63.9 vs 60.8. The BenchLeader Index combines every independent quality benchmark; Grok 4.6 (medium) is ahead overall as of 2026-09-10, but check the category scores for your use.
Which is cheaper, GPT-5 Pro or Grok 4.6?
Grok 4.6 is cheaper: $3.00 against $41.25 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Grok 4.6 accepts more context: 500k against 400k tokens.