BenchLeader

GPT-6 Astra vs Grok 4.5

Verdict
  • GPT-6 Astra (max) leads on quality: 70.6 vs 63.4.
  • GPT-6 Astra (max) is stronger in agents & tools, coding, composite, knowledge, maths, reasoning.
  • Grok 4.5 is stronger in human preference, long context, multimodal.
  • Grok 4.5 is 7.3× cheaper ($3.00 vs $22.00 per 1M blended).
MetricGPT-6 Astra (max)Grok 4.5
BenchLeader Index70.663.4
Agents & tools score68.258.6
Coding score79.059.9
Composite score73.166.0
Knowledge score79.876.0
Maths score77.0
Reasoning score70.968.9
Human preference score66.9
Long context score66.2
Multimodal score64.8
Blended price $/M$22.00$3.00
Output speed54 tok/s56 tok/s
Time to first answer328.6 s12.7 s
Context window1.1M500k
GPQA Diamond95.8%
FrontierMath Tiers 1–393.7%
FrontierMath Tier 497.6%
OTIS Mock AIME100.0%
SimpleQA Verified75.6%
Terminal-Bench58.2%
SimpleBench70.0%
SciCode56.5%
WeirdML46.4%
APEX-Agents34.2%
FrontierCode53.3%42.4%
Epoch Capabilities Index153.9
LMArena Text1471
LMArena Hard Prompts1495
LMArena Coding1523
LMArena WebDev17961556
LMArena Vision1291
LMArena Agent12.53.9
LiveBench82.2%75.8%
LiveBench Reasoning92.7%87.2%
LiveBench Coding80.4%68.6%
LiveBench Agentic Coding57.3%56.5%
LiveBench Mathematics96.8%90.8%
LiveBench Data Analysis83.0%73.0%
LiveBench Language89.4%82.8%
AA Intelligence Index39.1
AA-LCR79.3%
MMMU-Pro80.4%
AA-Omniscience25.3
GPQA Diamond (AA)93.1%
Humanity's Last Exam (AA)42.7%
SciCode (AA)55.0%
Terminal-Bench 2.1 (Vals)87.3%
Vals Index66.6
Kagi LLM Benchmark83.5%
ARC-AGI-197.5%
ARC-AGI-295.0%
ARC-AGI-362.7%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.