BenchLeader

GPT-6 Astra vs Grok 4.6

Verdict
  • GPT-6 Astra (max) leads on quality: 70.6 vs 63.8.
  • GPT-6 Astra (max) is stronger in agents & tools, coding, knowledge, maths, reasoning.
  • Grok 4.6 (medium) is stronger in composite, long context.
  • Grok 4.6 (medium) is 7.3× cheaper ($3.00 vs $22.00 per 1M blended).
MetricGPT-6 Astra (max)Grok 4.6 (medium)
BenchLeader Index70.663.8
Agents & tools score68.2
Coding score79.063.9
Composite score73.183.2
Knowledge score79.877.3
Maths score77.0
Reasoning score70.962.1
Long context score67.1
Blended price $/M$22.00$3.00
Output speed54 tok/s55 tok/s
Time to first answer328.6 s36.7 s
Context window1.1M500k
GPQA Diamond95.8%
FrontierMath Tiers 1–393.7%
FrontierMath Tier 497.6%
OTIS Mock AIME100.0%
SimpleQA Verified75.6%
Terminal-Bench58.2%
SciCode56.5%54.6%
FrontierCode53.3%
LMArena WebDev1796
LMArena Agent12.5
LiveBench82.2%
LiveBench Reasoning92.7%
LiveBench Coding80.4%
LiveBench Agentic Coding57.3%
LiveBench Mathematics96.8%
LiveBench Data Analysis83.0%
LiveBench Language89.4%
AA Intelligence Index43.0
AA-LCR81.0%
AA-Omniscience28
GPQA Diamond (AA)93.5%
Humanity's Last Exam (AA)42.1%
SciCode (AA)55.9%
Terminal-Bench 2.1 (Vals)87.3%
Vals Index66.6
ARC-AGI-197.5%87.5%
ARC-AGI-295.0%61.3%
ARC-AGI-362.7%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.