BenchLeader

GPT-5.5 vs Qwen3.5 397B A17B

Verdict
  • GPT-5.5 leads on quality: 66.7 vs 58.7.
  • GPT-5.5 is stronger in coding, composite, human preference, knowledge, long context, multimodal, reasoning.
  • Qwen3.5 397B A17B is stronger in agents & tools, instruction following, maths.
  • Qwen3.5 397B A17B is 29× cheaper ($0.387 vs $11.25 per 1M blended).
MetricGPT-5.5Qwen3.5 397B A17B
BenchLeader Index66.758.7
Agents & tools score68.569.4
Coding score57.251.7
Composite score77.853.1
Human preference score67.963.6
Instruction following score73.676.1
Knowledge score73.949.5
Long context score68.865.2
Multimodal score64.961.6
Reasoning score70.163.0
Maths score50.1
Blended price $/M$11.25$0.387
Output speed85 tok/s78 tok/s
Time to first answer58.6 s42.7 s
Context window1.1M262k
GPQA Diamond85.9%
FrontierMath Tiers 1–329.5%
OTIS Mock AIME88.9%
Terminal-Bench84.7%
SimpleBench69.0%
Remote Labor Index6.3%
APEX-Agents38.5%
FrontierCode43.0%
Epoch Capabilities Index159.1147.0
LMArena Text14771441
LMArena Hard Prompts14981463
LMArena Coding15091491
LMArena WebDev14581399
LMArena Vision12961265
LMArena Agent3
AA Intelligence Index38.619.1
IFBench75.8%78.8%
AA-LCR84.3%77.3%
MMMU-Pro79.9%77.3%
AA-Omniscience20.5-30.8
Terminal-Bench Hard60.6%40.9%
GPQA Diamond (AA)93.5%89.3%
Humanity's Last Exam (AA)45.8%29.0%
SciCode (AA)55.8%44.8%
τ²-Bench Telecom (AA)93.9%95.6%
AIME 202694.2%
HMMT February 202687.9%
HiL-Bench39.7%
EQ-Bench 41315
Kagi LLM Benchmark88.8%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

GPT-5.5 vs Qwen3.5 397B A17B: questions

Is GPT-5.5 better than Qwen3.5 397B A17B?
GPT-5.5 leads on quality: 66.7 vs 58.7. The BenchLeader Index combines every independent quality benchmark; GPT-5.5 is ahead overall as of 2026-09-10, but check the category scores for your use.
Is GPT-5.5 better than Qwen3.5 397B A17B for coding?
GPT-5.5 scores higher in coding (57 vs 52 on the category index, where 50 is average).
Is GPT-5.5 better than Qwen3.5 397B A17B for agentic tasks?
Qwen3.5 397B A17B scores higher in agentic tasks (69 vs 69 on the category index, where 50 is average).
Which is cheaper, GPT-5.5 or Qwen3.5 397B A17B?
Qwen3.5 397B A17B is cheaper: $0.387 against $11.25 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.5 or Qwen3.5 397B A17B?
GPT-5.5 streams faster: 85 against 78 output tokens per second.
Which has the larger context window?
GPT-5.5 accepts more context: 1.1M against 262k tokens.