BenchLeader

Claude Opus 5.5 vs GPT-5.6 Luna

Verdict
  • Claude Opus 5.5 (thinking) leads on quality: 70.7 vs 60.8.
  • Claude Opus 5.5 (thinking) is stronger in composite, knowledge, long context, multimodal, reasoning.
  • GPT-5.6 Luna (xhigh) is stronger in agents & tools, coding, human preference, maths.
  • GPT-5.6 Luna (xhigh) is 18× cheaper ($0.450 vs $8.00 per 1M blended).
  • GPT-5.6 Luna (xhigh) streams 3.0× faster (143 vs 48 tokens per second).
MetricClaude Opus 5.5 (thinking)GPT-5.6 Luna (xhigh)
BenchLeader Index70.760.8
Composite score95.071.3
Knowledge score85.058.3
Long context score68.567.0
Multimodal score71.961.4
Reasoning score95.059.5
Agents & tools score54.1
Coding score59.6
Human preference score64.8
Maths score67.0
Blended price $/M$8.00$0.450
Output speed48 tok/s143 tok/s
Time to first answer4.2 s40.2 s
Context window1M1.1M
SimpleBench46.8%
SciCode50.0%
LMArena Text1453
LMArena Hard Prompts1473
LMArena Coding1498
LMArena WebDev1519
LMArena Vision1259
LMArena Agent-0.4
AA Intelligence Index57.634.6
AA-LCR84.7%81.7%
MMMU-Pro87.7%78.5%
AA-Omniscience46.4-10.8
GPQA Diamond (AA)89.5%
Humanity's Last Exam (AA)61.4%37.0%
SciCode (AA)66.9%50.5%
ARC-AGI-187.7%
ARC-AGI-247.6%
ARC-AGI-30.0%
CritPt31.7%20.6%
GDPval (AA)67.3%44.2%
τ²-Bench Banking (AA)28.7%
LMArena Maths1477
LMArena Creative Writing1411
LMArena Instruction Following1442
LMArena Multi-turn1458
LMArena Longer Queries1450
LMArena Document1462
DeepSWE56.9%
LMCA48.5%
DTBench89.1%
CursorBench33.0%
Terminal-Bench 4.0 (AA)59.6%3.5%
Terminal-Bench 2.1 (AA)77.9%
AutomationBench69.5%
GDP.pdf26.2%
Harvey LAB91.2%
AA-Omniscience: accuracy66.2%42.5%
AA-Omniscience: non-hallucination41.4%7.5%
AA-Briefcase1822

Data as of 2026-09-22. Best configuration of each model; every score links to its source on the model pages.

Claude Opus 5.5 vs GPT-5.6 Luna: questions

Is Claude Opus 5.5 better than GPT-5.6 Luna?
Claude Opus 5.5 (thinking) leads on quality: 70.7 vs 60.8. The BenchLeader Index combines every independent quality benchmark; Claude Opus 5.5 (thinking) is ahead overall as of 2026-09-22, but check the category scores for your use.
Which is cheaper, Claude Opus 5.5 or GPT-5.6 Luna?
GPT-5.6 Luna is cheaper: $0.450 against $8.00 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Opus 5.5 or GPT-5.6 Luna?
GPT-5.6 Luna streams faster: 143 against 48 output tokens per second.
Which has the larger context window?
GPT-5.6 Luna accepts more context: 1.1M against 1M tokens.