BenchLeader

GPT-5.3-Codex vs Grok 4.3

Verdict
  • GPT-5.3-Codex leads on quality: 62.5 vs 60.7.
  • GPT-5.3-Codex is stronger in agents & tools, coding, composite, long context, multimodal.
  • Grok 4.3 (medium) is stronger in instruction following, knowledge.
  • Grok 4.3 (medium) is 3.1× cheaper ($1.56 vs $4.81 per 1M blended).
MetricGPT-5.3-CodexGrok 4.3 (medium)
BenchLeader Index62.560.7
Agents & tools score62.760.1
Coding score54.8
Composite score70.160.3
Instruction following score73.280.0
Knowledge score69.372.1
Long context score68.364.0
Multimodal score63.060.2
Blended price $/M$4.81$1.56
Output speed121 tok/s112 tok/s
Time to first answer58.2 s12.1 s
Context window400k1M
Terminal-Bench78.4%
WeirdML79.3%
APEX-Agents31.8%
Epoch Capabilities Index156.6
LMArena WebDev1409
AA Intelligence Index32.524.8
IFBench75.4%83.3%
AA-LCR83.3%75.0%
MMMU-Pro78.5%75.8%
AA-Omniscience10.916.7
Terminal-Bench Hard53.0%30.3%
GPQA Diamond (AA)91.5%89.0%
Humanity's Last Exam (AA)42.5%30.0%
τ²-Bench Telecom (AA)86.0%91.2%
HiL-Bench4.3%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

GPT-5.3-Codex vs Grok 4.3: questions

Is GPT-5.3-Codex better than Grok 4.3?
GPT-5.3-Codex leads on quality: 62.5 vs 60.7. The BenchLeader Index combines every independent quality benchmark; GPT-5.3-Codex is ahead overall as of 2026-09-10, but check the category scores for your use.
Is GPT-5.3-Codex better than Grok 4.3 for agentic tasks?
GPT-5.3-Codex scores higher in agentic tasks (63 vs 60 on the category index, where 50 is average).
Which is cheaper, GPT-5.3-Codex or Grok 4.3?
Grok 4.3 is cheaper: $1.56 against $4.81 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.3-Codex or Grok 4.3?
GPT-5.3-Codex streams faster: 121 against 112 output tokens per second.
Which has the larger context window?
Grok 4.3 accepts more context: 1M against 400k tokens.