BenchLeader

GPT-5.5 Pro vs GPT-5.6 Luna

Verdict
  • GPT-5.5 Pro (xhigh) leads on quality: 67.1 vs 61.1.
  • GPT-5.5 Pro (xhigh) is stronger in maths, reasoning.
  • GPT-5.6 Luna (xhigh) is stronger in agents & tools, coding, composite, human preference, knowledge, long context, multimodal.
  • GPT-5.6 Luna (xhigh) is 150× cheaper ($0.450 vs $67.50 per 1M blended).
  • GPT-5.6 Luna (xhigh) streams 31.0× faster (124 vs 4 tokens per second).
MetricGPT-5.5 Pro (xhigh)GPT-5.6 Luna (xhigh)
BenchLeader Index67.161.1
Maths score74.567.0
Reasoning score80.160.1
Agents & tools score54.1
Coding score59.8
Composite score73.1
Human preference score64.8
Knowledge score59.1
Long context score67.3
Multimodal score61.6
Blended price $/M$67.50$0.450
Output speed4 tok/s124 tok/s
Time to first answer3.5 s34.9 s
Context window1.1M1.1M
FrontierMath Tiers 1–387.7%
FrontierMath Tier 478.0%
SimpleBench46.8%
SciCode50.0%
LMArena Text1453
LMArena Hard Prompts1473
LMArena Coding1498
LMArena WebDev1519
LMArena Vision1259
LMArena Agent-0.4
AA Intelligence Index34.8
AA-LCR81.7%
MMMU-Pro78.5%
AA-Omniscience-10.8
GPQA Diamond (AA)89.5%
Humanity's Last Exam (AA)37.0%
SciCode (AA)50.5%
ARC-AGI-195.0%87.7%
ARC-AGI-284.2%47.6%
ARC-AGI-30.0%
CritPt30.6%20.6%
GDPval (AA)46.4%
τ²-Bench Banking (AA)28.7%
LMArena Maths1477
LMArena Creative Writing1411
LMArena Instruction Following1442
LMArena Multi-turn1458
LMArena Longer Queries1450
LMArena Document1462
DeepSWE56.9%
LMCA53.9%
DTBench96.0%
CursorBench57.7%

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

GPT-5.5 Pro vs GPT-5.6 Luna: questions

Is GPT-5.5 Pro better than GPT-5.6 Luna?
GPT-5.5 Pro (xhigh) leads on quality: 67.1 vs 61.1. The BenchLeader Index combines every independent quality benchmark; GPT-5.5 Pro (xhigh) is ahead overall as of 2026-09-19, but check the category scores for your use.
Which is cheaper, GPT-5.5 Pro or GPT-5.6 Luna?
GPT-5.6 Luna is cheaper: $0.450 against $67.50 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.5 Pro or GPT-5.6 Luna?
GPT-5.6 Luna streams faster: 124 against 4 output tokens per second.
Which has the larger context window?
Both accept 1.1M tokens of context.