BenchLeader

Claude Opus 5.5 vs Grok 4.7

Verdict
  • Claude Opus 5.5 (thinking) leads on quality: 70.7 vs 61.9.
  • Claude Opus 5.5 (thinking) is stronger in composite, knowledge, long context, multimodal, reasoning.
  • Grok 4.7 (xhigh) is stronger in agents & tools, coding.
  • Grok 4.7 (xhigh) is 2.7× cheaper ($3.00 vs $8.00 per 1M blended).
MetricClaude Opus 5.5 (thinking)Grok 4.7 (xhigh)
BenchLeader Index70.761.9
Composite score95.072.1
Knowledge score85.071.9
Long context score68.564.4
Multimodal score71.9
Reasoning score95.074.2
Agents & tools score50.8
Coding score59.6
Blended price $/M$8.00$3.00
Output speed48 tok/s39 tok/s
Time to first answer4.2 s0.8 s
Context window1M500k
Terminal-Bench37.6%
SciCode57.4%
LiveBench77.4%
LiveBench Reasoning82.7%
LiveBench Coding77.2%
LiveBench Agentic Coding54.0%
LiveBench Mathematics95.7%
LiveBench Data Analysis76.9%
LiveBench Language80.1%
LiveBench Instruction Following75.3%
AA Intelligence Index57.646.5
AA-LCR84.7%76.7%
MMMU-Pro87.7%
AA-Omniscience46.432.0
Humanity's Last Exam (AA)61.4%43.1%
SciCode (AA)66.9%57.4%
IOI57.7%
LegalBench84.4%
Terminal-Bench 2.1 (Vals)73.4%
Vals Index60.2
CritPt31.7%17.7%
GDPval (AA)67.3%59.8%
Code Migration44.8%
Excel Modeling Benchmark67.0%
Finance Agent v252.3%
Harvey's Legal Agent Benchmark12.1%
Legal Research Bench47.1%
MedCode49.5%
MedScribe89.4%
MysteryMechanism25.2%
Public Benefits Bench65.6%
SAGE40.8%
Tax Agent Bench65.6%
Vibe Code Bench v1.186.2%
CursorBench46.3%
Terminal-Bench 4.0 (AA)59.6%25.8%
AutomationBench69.5%65.6%
GDP.pdf26.2%20.0%
MLCR15.0%
Harvey LAB91.2%
AA-Omniscience: accuracy66.2%47.5%
AA-Omniscience: non-hallucination41.4%70.7%
AA-Briefcase18221657

Data as of 2026-09-22. Best configuration of each model; every score links to its source on the model pages.

Claude Opus 5.5 vs Grok 4.7: questions

Is Claude Opus 5.5 better than Grok 4.7?
Claude Opus 5.5 (thinking) leads on quality: 70.7 vs 61.9. The BenchLeader Index combines every independent quality benchmark; Claude Opus 5.5 (thinking) is ahead overall as of 2026-09-22, but check the category scores for your use.
Which is cheaper, Claude Opus 5.5 or Grok 4.7?
Grok 4.7 is cheaper: $3.00 against $8.00 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Opus 5.5 or Grok 4.7?
Claude Opus 5.5 streams faster: 48 against 39 output tokens per second.
Which has the larger context window?
Claude Opus 5.5 accepts more context: 1M against 500k tokens.