BenchLeader

Grok 4.6 vs o3

Verdict
  • Grok 4.6 (medium) leads on quality: 63.9 vs 61.6.
  • Grok 4.6 (medium) is stronger in coding, composite, knowledge, long context, reasoning.
  • o3 is stronger in agents & tools, human preference, instruction following, maths, multimodal.
  • Grok 4.6 (medium) is 1.2× cheaper ($3.00 vs $3.50 per 1M blended).
  • o3 streams 2.1× faster (107 vs 52 tokens per second).
MetricGrok 4.6 (medium)o3
BenchLeader Index63.961.6
Coding score64.162.0
Composite score83.554.5
Knowledge score77.658.0
Long context score67.163.8
Reasoning score62.161.5
Agents & tools score70.3
Human preference score62.3
Instruction following score65.0
Maths score78.6
Multimodal score57.9
Blended price $/M$3.00$3.50
Output speed52 tok/s107 tok/s
Time to first answer26.1 s6.3 s
Context window500k200k
SciCode54.6%
Epoch Capabilities Index146.9
LMArena Text1432
LMArena Hard Prompts1441
LMArena Coding1460
LMArena Vision1214
AA Intelligence Index43.020.2
IFBench71.4%
AA-LCR81.0%74.7%
MMMU-Pro70.1%
AA-Omniscience28-15.6
Terminal-Bench Hard37.1%
GPQA Diamond (AA)93.5%82.7%
Humanity's Last Exam (AA)42.1%20.1%
SciCode (AA)55.9%
τ²-Bench Telecom (AA)80.7%
PRBench Finance47.7%
PRBench Legal48.6%
MMMU (validation)82.9%
MMMU-Pro (official)76.4%
Kagi LLM Benchmark67.6%
IFEval (HELM)86.9%
Omni-MATH (HELM)71.4%
WildBench (HELM)86.1%
MMLU-Pro (HELM)85.9%
GPQA Diamond (HELM)75.3%
HELM Capabilities mean81.1%
Aider Polyglot81.3%
SWE-bench Verified (bash only)58.4%
SWE-bench Verified (any scaffold)58.4%
ARC-AGI-187.5%
ARC-AGI-261.3%
BFCL Overall63.0%

Data as of 2026-09-12. Best configuration of each model; every score links to its source on the model pages.

Grok 4.6 vs o3: questions

Is Grok 4.6 better than o3?
Grok 4.6 (medium) leads on quality: 63.9 vs 61.6. The BenchLeader Index combines every independent quality benchmark; Grok 4.6 (medium) is ahead overall as of 2026-09-12, but check the category scores for your use.
Is Grok 4.6 better than o3 for coding?
Grok 4.6 scores higher in coding (64 vs 62 on the category index, where 50 is average).
Which is cheaper, Grok 4.6 or o3?
Grok 4.6 is cheaper: $3.00 against $3.50 per million tokens, blended at three input tokens per output token.
Which is faster, Grok 4.6 or o3?
o3 streams faster: 107 against 52 output tokens per second.
Which has the larger context window?
Grok 4.6 accepts more context: 500k against 200k tokens.