BenchLeader

GPT-5.6 Sol vs Grok 4.6

Verdict
  • GPT-5.6 Sol (high) leads on quality: 68.0 vs 63.8.
  • GPT-5.6 Sol (high) is stronger in agents & tools, coding, instruction following, long context, multimodal, reasoning.
  • Grok 4.6 (medium) is stronger in composite, knowledge.
  • Grok 4.6 (medium) is 2.7× cheaper ($3.00 vs $8.00 per 1M blended).
MetricGPT-5.6 Sol (high)Grok 4.6 (medium)
BenchLeader Index68.063.8
Agents & tools score87.6
Coding score72.063.9
Composite score82.583.2
Instruction following score67.6
Knowledge score73.777.3
Long context score67.467.1
Multimodal score66.5
Reasoning score62.962.1
Blended price $/M$8.00$3.00
Output speed59 tok/s55 tok/s
Time to first answer31.4 s36.7 s
Context window1M500k
SciCode56.9%54.6%
WeirdML88.8%
AA Intelligence Index42.543.0
IFBench69.2%
AA-LCR81.7%81.0%
MMMU-Pro81.8%
AA-Omniscience20.428
Terminal-Bench Hard62.1%
GPQA Diamond (AA)92.8%93.5%
Humanity's Last Exam (AA)46.0%42.1%
SciCode (AA)57.8%55.9%
τ²-Bench Telecom (AA)83.3%
EnigmaEval37.1%
ARC-AGI-197.0%87.5%
ARC-AGI-285.4%61.3%
ARC-AGI-32.2%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.