BenchLeader

GPT-5 vs Grok 4

Verdict
  • GPT-5 and Grok 4 are level on quality (59.7 vs 59.8).
  • GPT-5 is stronger in composite, human preference, knowledge, maths, multimodal, reasoning.
  • Grok 4 is stronger in agents & tools, coding, instruction following, long context.
  • Grok 4 is 2.2× cheaper ($1.56 vs $3.44 per 1M blended).
MetricGPT-5Grok 4
BenchLeader Index59.759.8
Agents & tools score50.053.4
Coding score60.961.2
Composite score52.8
Human preference score61.759.8
Instruction following score64.067.7
Knowledge score60.856.0
Long context score63.077.7
Maths score72.858.0
Multimodal score58.954.5
Reasoning score64.258.3
Blended price $/M$3.44$1.56
Output speed106 tok/s
Time to first answer54.4 s
Context window400k256k
GPQA Diamond87.0%
OTIS Mock AIME84.0%
Terminal-Bench49.6%27.2%
SimpleBench60.5%
Fiction.LiveBench 120k96.9%
SciCode42.9%
Cybench43.0%
Remote Labor Index1.7%
WeirdML39.8%45.7%
APEX-Agents18.3%15.2%
Epoch Capabilities Index150146.4
LMArena Text14271411
LMArena Hard Prompts14491420
LMArena Coding14621435
LMArena Vision12321210
AA Intelligence Index23.0
IFBench73.1%
AA-LCR78.2%
MMMU-Pro74.2%
AA-Omniscience-8.7
Terminal-Bench Hard32.6%
GPQA Diamond (AA)85.3%
Humanity's Last Exam (AA)28.5%
τ²-Bench Telecom (AA)84.8%
AIME (Vals)90.6%
LiveCodeBench83.3%
MMLU-Pro85.3%
LegalBench83.2%
CorpFin66.0%
TaxEval65.1%
MedQA92.5%
MGSM90.9%
SWE-bench (Vals)57.8%
GPQA Diamond (Vals)88.1%
IMO 202521.4%
MathArena Apex2.1%
PRBench Finance51.3%
PRBench Legal49.0%
VISTA49.7%
MultiNRC52.1%
TutorBench55.3%
Kagi LLM Benchmark72.7%73.6%
IFEval (HELM)87.5%94.9%
Omni-MATH (HELM)64.7%60.3%
WildBench (HELM)85.7%79.7%
MMLU-Pro (HELM)86.3%85.1%
GPQA Diamond (HELM)79.1%72.6%
HELM Capabilities mean80.7%78.5%
Aider Polyglot88.0%79.6%
SWE-bench Verified (any scaffold)75.6%
ARC-AGI-179.6%
ARC-AGI-229.4%
BFCL Overall63.0%

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

GPT-5 vs Grok 4: questions

Is GPT-5 better than Grok 4?
GPT-5 and Grok 4 are level on quality (59.7 vs 59.8). The BenchLeader Index combines every independent quality benchmark; Grok 4 is ahead overall as of 2026-09-13, but check the category scores for your use.
Is GPT-5 better than Grok 4 for coding?
Grok 4 scores higher in coding (61 vs 61 on the category index, where 50 is average).
Is GPT-5 better than Grok 4 for agentic tasks?
Grok 4 scores higher in agentic tasks (53 vs 50 on the category index, where 50 is average).
Which is cheaper, GPT-5 or Grok 4?
Grok 4 is cheaper: $1.56 against $3.44 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
GPT-5 accepts more context: 400k against 256k tokens.