BenchLeader

GPT-5.5 Pro vs o3

Verdict
  • GPT-5.5 Pro (xhigh) leads on quality: 67.1 vs 60.3.
  • GPT-5.5 Pro (xhigh) is stronger in maths, reasoning.
  • o3 is stronger in agents & tools, coding, composite, human preference, instruction following, knowledge, long context, multimodal.
  • o3 is 19× cheaper ($3.50 vs $67.50 per 1M blended).
  • o3 streams 31.8× faster (127 vs 4 tokens per second).
MetricGPT-5.5 Pro (xhigh)o3
BenchLeader Index67.160.3
Maths score74.571.2
Reasoning score80.152.9
Agents & tools score70.2
Coding score61.9
Composite score54.6
Human preference score62.2
Instruction following score65.1
Knowledge score57.2
Long context score63.7
Multimodal score57.7
Blended price $/M$67.50$3.50
Output speed4 tok/s127 tok/s
Time to first answer3.5 s5.9 s
Context window1.1M200k
FrontierMath Tiers 1–387.7%
FrontierMath Tier 478.0%
Epoch Capabilities Index146.9
LMArena Text1432
LMArena Hard Prompts1441
LMArena Coding1460
LMArena Vision1214
AA Intelligence Index20.2
IFBench71.4%
AA-LCR74.7%
MMMU-Pro70.1%
AA-Omniscience-15.6
Terminal-Bench Hard37.1%
GPQA Diamond (AA)82.7%
Humanity's Last Exam (AA)20.1%
τ²-Bench Telecom (AA)80.7%
PRBench Finance47.7%
PRBench Legal48.6%
MMMU (validation)82.9%
MMMU-Pro (official)76.4%
Kagi LLM Benchmark67.6%
IFEval (HELM)86.9%
Omni-MATH (HELM)71.4%
WildBench (HELM)86.1%
MMLU-Pro (HELM)85.9%
GPQA Diamond (HELM)75.3%
HELM Capabilities mean81.1%
Aider Polyglot81.3%
SWE-bench Verified (bash only)58.4%
SWE-bench Verified (any scaffold)58.4%
ARC-AGI-195.0%
ARC-AGI-284.2%
BFCL Overall63.0%
CritPt30.6%1.1%
FORTRESS16.0%
PropensityBench10.5%
SciPredict17.9%
VTB13.7%
LMArena Maths1448
LMArena Creative Writing1382
LMArena Instruction Following1402
LMArena Multi-turn1420
LMArena Longer Queries1410
METR Time Horizons63.6%
LMCA53.9%
DTBench96.0%
ForecastBench62.5%

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

GPT-5.5 Pro vs o3: questions

Is GPT-5.5 Pro better than o3?
GPT-5.5 Pro (xhigh) leads on quality: 67.1 vs 60.3. The BenchLeader Index combines every independent quality benchmark; GPT-5.5 Pro (xhigh) is ahead overall as of 2026-09-19, but check the category scores for your use.
Which is cheaper, GPT-5.5 Pro or o3?
o3 is cheaper: $3.50 against $67.50 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.5 Pro or o3?
o3 streams faster: 127 against 4 output tokens per second.
Which has the larger context window?
GPT-5.5 Pro accepts more context: 1.1M against 200k tokens.