BenchLeader

GPT-6 Sol vs Grok 4.7

Verdict
  • GPT-6 Sol (max) leads on quality: 67.7 vs 61.6.
  • GPT-6 Sol (max) is stronger in composite, knowledge, long context, multimodal, reasoning.
  • Grok 4.7 (xhigh) is stronger in agents & tools, coding.
  • Grok 4.7 (xhigh) is 1.3× cheaper ($3.00 vs $4.00 per 1M blended).
  • GPT-6 Sol (max) streams 3.2× faster (126 vs 39 tokens per second).
MetricGPT-6 Sol (max)Grok 4.7 (xhigh)
BenchLeader Index67.761.6
Composite score86.271.7
Knowledge score75.471.6
Long context score67.764.1
Multimodal score67.1
Reasoning score95.073.0
Agents & tools score50.8
Coding score59.6
Blended price $/M$4.00$3.00
Output speed126 tok/s39 tok/s
Time to first answer107.2 s0.9 s
Context window1.1M500k
Terminal-Bench37.6%
SciCode57.4%
LiveBench77.4%
LiveBench Reasoning82.7%
LiveBench Coding77.2%
LiveBench Agentic Coding54.0%
LiveBench Mathematics95.7%
LiveBench Data Analysis76.9%
LiveBench Language80.1%
LiveBench Instruction Following75.3%
AA Intelligence Index47.546.5
AA-LCR83.7%76.7%
MMMU-Pro83.3%
AA-Omniscience27.132.0
Humanity's Last Exam (AA)47.9%43.1%
SciCode (AA)57.6%57.4%
IOI57.7%
LegalBench84.4%
Terminal-Bench 2.1 (Vals)73.4%
Vals Index60.2
CritPt30.9%17.7%
GDPval (AA)49.4%59.8%
Code Migration44.8%
Excel Modeling Benchmark67.0%
Finance Agent v252.3%
Harvey's Legal Agent Benchmark12.1%
Legal Research Bench47.1%
MedCode49.5%
MedScribe89.4%
MysteryMechanism25.2%
Public Benefits Bench65.6%
SAGE40.8%
Tax Agent Bench65.6%
Vibe Code Bench v1.186.2%
CursorBench46.3%
Terminal-Bench 4.0 (AA)43.9%25.8%
AutomationBench61.6%65.6%
GDP.pdf24.8%20.0%
MLCR16.1%15.0%
AA-Omniscience: accuracy54.5%47.5%
AA-Omniscience: non-hallucination39.9%70.7%
AA-Briefcase14831657

Data as of 2026-09-23. Best configuration of each model; every score links to its source on the model pages.

GPT-6 Sol vs Grok 4.7: questions

Is GPT-6 Sol better than Grok 4.7?
GPT-6 Sol (max) leads on quality: 67.7 vs 61.6. The BenchLeader Index combines every independent quality benchmark; GPT-6 Sol (max) is ahead overall as of 2026-09-23, but check the category scores for your use.
Which is cheaper, GPT-6 Sol or Grok 4.7?
Grok 4.7 is cheaper: $3.00 against $4.00 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-6 Sol or Grok 4.7?
GPT-6 Sol streams faster: 126 against 39 output tokens per second.
Which has the larger context window?
GPT-6 Sol accepts more context: 1.1M against 500k tokens.