BenchLeader

Claude Opus 4.6 vs GPT-5.5

Verdict
  • GPT-5.5 leads on quality: 66.6 vs 63.1.
  • Claude Opus 4.6 is stronger in agents & tools, coding, maths.
  • GPT-5.5 is stronger in composite, human preference, instruction following, knowledge, long context, multimodal, reasoning.
  • They cost about the same ($10.00 per 1M blended).
  • GPT-5.5 streams 2.3× faster (88 vs 38 tokens per second).
MetricClaude Opus 4.6GPT-5.5
BenchLeader Index63.166.6
Agents & tools score68.868.4
Coding score62.357.3
Composite score69.277.7
Human preference score64.667.6
Instruction following score53.773.3
Knowledge score70.573.7
Long context score65.568.8
Maths score65.6
Multimodal score63.664.9
Reasoning score62.369.7
Blended price $/M$10.00$11.25
Output speed38 tok/s88 tok/s
Time to first answer2.0 s62.3 s
Context window1M1.1M
GPQA Diamond90.5%
OTIS Mock AIME94.4%
SWE-bench Verified (Epoch)78.7%
Humanity's Last Exam19.0%
Terminal-Bench79.8%84.7%
SimpleBench67.6%69.0%
Cybench93.0%
Remote Labor Index4.2%6.3%
WeirdML77.9%
APEX-Agents32.4%38.5%
FrontierCode26.6%43.0%
GSO-Bench33.3%
Epoch Capabilities Index155.3159.1
LMArena Text14981477
LMArena Hard Prompts15271498
LMArena Coding15461509
LMArena WebDev15371458
LMArena Vision13111296
LMArena Agent3
AA Intelligence Index31.938.6
IFBench53.1%75.8%
AA-LCR78.0%84.3%
MMMU-Pro75.4%79.9%
AA-Omniscience13.720.5
Terminal-Bench Hard48.5%60.6%
GPQA Diamond (AA)89.6%93.5%
Humanity's Last Exam (AA)39.9%45.8%
SciCode (AA)55.8%
τ²-Bench Telecom (AA)92.1%93.9%
HiL-Bench38.3%39.7%
EQ-Bench 412231315
Kagi LLM Benchmark72.4%88.8%
SWE-bench Verified (bash only)75.6%
SWE-bench Verified (any scaffold)75.6%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.