BenchLeader
OpenAIReasoning modelAuto-detected

GPT-5.5

Best configuration ranks #17 of 610 on the BenchLeader Index at 66.6 ±3.9. Last measured 8 Sept 2026. Released 23 Apr 2026.

Blended price
$11.25/M
$5.00 in · $30.00 out
Output speed
88 tok/s
First answer
62 s
first token 3.50 s
Context
1.1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index67
  2. Reasoning70
  3. Coding57
  4. Agents & tools68
  5. Knowledge74
  6. Instruction following73
  7. Human preference68
  8. Multimodal65
  9. Long context69
  10. Composite78

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning56.2#16079 tok/s0.99 s$0.011Agents & tools 76 · Coding 58 · Composite 58 · Instruction following 48 · Knowledge 62 · Long context 58 · Maths 49 · Multimodal 56 · Reasoning 57
low62.3#5775 tok/s1.65 s$0.011Agents & tools 79 · Coding 60 · Composite 68 · Instruction following 63 · Knowledge 71 · Long context 67 · Maths 61 · Multimodal 64 · Reasoning 57
medium64.9#3082 tok/s7.61 s$0.011Agents & tools 84 · Coding 62 · Composite 72 · Instruction following 69 · Knowledge 73 · Long context 68 · Multimodal 66 · Reasoning 66
high66.3#1978 tok/s14 s$0.011Agents & tools 75 · Coding 63 · Composite 76 · Human preference 68 · Instruction following 70 · Knowledge 73 · Long context 69 · Multimodal 65 · Reasoning 64
xhigh63.7#4388 tok/s62 s$0.012Agents & tools 56 · Coding 65 · Composite 67 · Knowledge 64 · Maths 71 · Reasoning 70
defaultbest66.6#1788 tok/s62 s$0.011Agents & tools 68 · Coding 57 · Composite 78 · Human preference 68 · Instruction following 73 · Knowledge 74 · Long context 69 · Multimodal 65 · Reasoning 70

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
GPQA Diamond77.3%#12490.7%#34Epoch AI Benchmarking Hub
SimpleBench69.0%#15SimpleBench
LMArena Hard Prompts1502#231498#26LMArena
LiveBench Reasoningnot in index89.7%#9LiveBench
GPQA Diamond (AA)not in index76.8%#22091.0%#5492.6%#3293.2%#2293.5%#14Artificial Analysis
Humanity's Last Exam (AA)not in index13.7%#20732.7%#9542.4%#4045.0%#2945.8%#27Artificial Analysis
GPQA Diamond (Vals)not in index93.2%#10Vals AI
Kagi LLM Benchmark88.8%#2Kagi LLM Benchmark
ARC-AGI-176.2%#9092.2%#3994.5%#2495.0%#22ARC Prize
ARC-AGI-233.3%#8470.4%#3783.3%#2485.0%#18ARC Prize
ARC-AGI-30.4%#24ARC Prize

Coding

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
SciCode47.3%#7451.6%#4953.5%#4155.9%#2056.1%#17SciCode
WeirdML67.2%#3583.9%#1384.9%#12WeirdML
FrontierCode43.0%#10Cognition
GSO-Bench40.2%#5GSO-Bench
LMArena Coding1520#281509#48LMArena
LMArena WebDev1487#431510#371458#52LMArena
LiveBench Codingnot in index82.2%#6LiveBench
SciCode (AA)not in index54.5%#3756.1%#2555.8%#27Artificial Analysis
LiveCodeBench85.3%#33Vals AI
SWE-bench (Vals)not in index82.6%#19Vals AI

Agents & tools

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
Terminal-Bench84.7%#1Terminal-Bench
OSWorld-Verified 2.013.0%#9OSWorld
Remote Labor Index6.3%#3Scale AI / CAIS
APEX-Agents38.5%#1238.5%#12Mercor
LMArena Agent5.3#103#17LMArena
LiveBench Agentic Codingnot in index54.0%#25LiveBench
Terminal-Bench Hard49.2%#2452.3%#1957.6%#1159.9%#960.6%#7Artificial Analysis
τ²-Bench Telecom (AA)not in index69.3%#15383.9%#10291.8%#5693.0%#4793.9%#39Artificial Analysis
Terminal-Bench 2.1 (Vals)76.4%#13Vals AI
MCP Atlas75.3%#16Scale AI SEAL
HiL-Bench39.7%#7Scale AI SEAL

Maths

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
FrontierMath Tiers 1–385.3%#8Epoch AI Benchmarking Hub
FrontierMath Tier 472.5%#14Epoch AI Benchmarking Hub
OTIS Mock AIME57.8%#15084.4%#83Epoch AI Benchmarking Hub
ProofBench50.0%#17Vals AI
LiveBench Mathematicsnot in index95.9%#6LiveBench
AIME 2026100.0%#1MathArena
HMMT February 202698.5%#1MathArena
MathArena Apex80.2%#2MathArena

Knowledge

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
SimpleQA Verified63.0%#11Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index81.6%#2LiveBench
AA-Omniscience-4.8#11615.1#5218.1#4718.8#4620.5#41Artificial Analysis
MMLU-Pro88.1%#21Vals AI
LegalBench86.5%#11Vals AI
CorpFin68.4%#8Vals AI
TaxEval75.0%#24Vals AI

Instruction following

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
LiveBench Languagenot in index87.4%#7LiveBench
IFBench46.1%#19364.3%#10871.0%#6371.6%#5275.8%#26Artificial Analysis

Human preference

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
LMArena Text1482#191477#24LMArena
EQ-Bench 41315#4EQ-Bench

Multimodal

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
LMArena Vision1294#181296#15LMArena
MMMU-Pro71.4%#10779.0%#4581.2%#2781.1%#2879.9%#38Artificial Analysis

Long context

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
AA-LCR64.0%#19981.0%#3783.0%#1284.3%#484.3%#4Artificial Analysis

Composite

Benchmarkno reasoninglowmediumhighxhighdefaultSourceTrend
Epoch Capabilities Indexnot in index159.1#8Epoch AI Benchmarking Hub
LiveBench80.2%#6LiveBench
AA Intelligence Index23.2#13530.7#7534.1#6037.3#4838.6#44Artificial Analysis
Vals Indexnot in index57.4#16Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Amazon Bedrock (US)88 tok/s1.62 s$5.50$33.001.1M
OpenAI Fast80 tok/s2.03 s$12.50$75.001.1M
OpenAI50 tok/s2.56 s$5.00$30.001.1M
Azure36 tok/s6.61 s$5.00$30.001.1M
Azure (EU)31 tok/s4.44 s$5.50$33.001.1M
Azure (US)28 tok/s16 s$5.50$33.001.1M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.500 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.011$0.00971.1 min
Summarise a 30-page report12,000 / 600$0.078$0.0371.2 min
Code edit6,000 / 1,500$0.075$0.0551.3 min
Agentic coding session60,000 / 4,000$0.420$0.2171.8 min
Structured extraction2,000 / 200$0.016$0.00931.1 min

See also

Data as of 9 Sept 2026. Compare these configurations.