BenchLeader

Claude Opus 4.7 vs GPT-5.6 Terra

Verdict
  • Claude Opus 4.7 leads on quality: 65.5 vs 64.2.
  • Claude Opus 4.7 is stronger in coding, composite, human preference, knowledge, multimodal, reasoning.
  • GPT-5.6 Terra (xhigh) is stronger in agents & tools, instruction following, long context, maths.
  • GPT-5.6 Terra (xhigh) is 2.2× cheaper ($4.50 vs $10.00 per 1M blended).
  • GPT-5.6 Terra (xhigh) streams 2.0× faster (93 vs 46 tokens per second).
MetricClaude Opus 4.7GPT-5.6 Terra (xhigh)
BenchLeader Index65.564.2
Agents & tools score69.469.6
Coding score64.863.7
Composite score80.377.1
Human preference score68.866.3
Instruction following score58.565.0
Knowledge score66.460.9
Long context score65.966.0
Maths score65.271.0
Multimodal score65.763.0
Reasoning score66.358.2
Blended price $/M$10.00$4.50
Output speed46 tok/s93 tok/s
Time to first answer21.2 s32.5 s
Context window1M1M
Humanity's Last Exam36.2%
Terminal-Bench80.2%
SimpleBench61.7%48.9%
SciCode51.6%
WeirdML76.4%
FrontierCode38.5%
ProofBench74.0%
GSO-Bench44.1%
Epoch Capabilities Index156.3
LMArena Text14951466
LMArena Hard Prompts15191491
LMArena Coding15471519
LMArena WebDev15571521
LMArena Vision13171269
LMArena Agent1.5
AA Intelligence Index40.738.2
IFBench58.6%66.3%
AA-LCR78.7%79.0%
MMMU-Pro78.8%79.5%
AA-Omniscience27.3-3.0
Terminal-Bench Hard51.5%62.9%
GPQA Diamond (AA)91.4%90.8%
Humanity's Last Exam (AA)42.3%41.9%
SciCode (AA)52.3%
τ²-Bench Telecom (AA)88.6%80.4%
AIME (Vals)96.3%
LiveCodeBench85.1%85.9%
MMLU-Pro89.9%86.7%
IOI47.1%65.3%
LegalBench85.3%
CorpFin66.1%65.3%
TaxEval75.3%76.2%
Terminal-Bench 2.1 (Vals)68.5%
SWE-bench (Vals)82.0%
GPQA Diamond (Vals)90.2%90.9%
Vals Index56.1
HiL-Bench41.7%
EQ-Bench 41311
Kagi LLM Benchmark73.3%
ARC-AGI-194.0%
ARC-AGI-274.2%
ARC-AGI-30.7%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.