BenchLeader

GPT-5.5 Instant vs o3

Verdict
  • o3 leads on quality: 61.5 vs 60.5.
  • GPT-5.5 Instant is stronger in agents & tools, composite, human preference, instruction following, knowledge, multimodal, reasoning.
  • o3 is stronger in coding, long context, maths.
  • o3 is 3.2× cheaper ($3.50 vs $11.25 per 1M blended).
MetricGPT-5.5 Instanto3
BenchLeader Index60.561.5
Agents & tools score70.770.3
Coding score61.061.9
Composite score63.054.6
Human preference score67.462.2
Instruction following score69.865.0
Knowledge score69.657.2
Long context score61.363.8
Maths score42.578.6
Multimodal score59.357.7
Reasoning score62.561.5
Blended price $/M$11.25$3.50
Output speed131 tok/s144 tok/s
Time to first answer16.4 s4.1 s
Context window400k200k
GPQA Diamond82.5%
FrontierMath Tiers 1–326.3%
FrontierMath Tier 42.4%
OTIS Mock AIME68.1%
SciCode48.6%
Epoch Capabilities Index142.5146.9
LMArena Text14741432
LMArena Hard Prompts14911441
LMArena Coding15141460
LMArena Vision12511214
AA Intelligence Index26.820.2
IFBench71.5%71.4%
AA-LCR70.0%74.7%
MMMU-Pro70.1%
AA-Omniscience11.2-15.6
Terminal-Bench Hard42.4%37.1%
GPQA Diamond (AA)84.7%82.7%
Humanity's Last Exam (AA)21.6%20.1%
SciCode (AA)52.5%
τ²-Bench Telecom (AA)49.4%80.7%
PRBench Finance47.7%
PRBench Legal48.6%
MMMU (validation)82.9%
MMMU-Pro (official)76.4%
Kagi LLM Benchmark67.6%
IFEval (HELM)86.9%
Omni-MATH (HELM)71.4%
WildBench (HELM)86.1%
MMLU-Pro (HELM)85.9%
GPQA Diamond (HELM)75.3%
HELM Capabilities mean81.1%
Aider Polyglot81.3%
SWE-bench Verified (bash only)58.4%
SWE-bench Verified (any scaffold)58.4%
BFCL Overall63.0%

Data as of 2026-09-15. Best configuration of each model; every score links to its source on the model pages.

GPT-5.5 Instant vs o3: questions

Is GPT-5.5 Instant better than o3?
o3 leads on quality: 61.5 vs 60.5. The BenchLeader Index combines every independent quality benchmark; o3 is ahead overall as of 2026-09-15, but check the category scores for your use.
Is GPT-5.5 Instant better than o3 for coding?
o3 scores higher in coding (62 vs 61 on the category index, where 50 is average).
Is GPT-5.5 Instant better than o3 for agentic tasks?
GPT-5.5 Instant scores higher in agentic tasks (71 vs 70 on the category index, where 50 is average).
Which is cheaper, GPT-5.5 Instant or o3?
o3 is cheaper: $3.50 against $11.25 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.5 Instant or o3?
o3 streams faster: 144 against 131 output tokens per second.
Which has the larger context window?
GPT-5.5 Instant accepts more context: 400k against 200k tokens.