BenchLeader

Claude Haiku 5.5 vs GPT-5.6 Sol

Verdict
  • GPT-5.6 Sol (max) leads on quality: 67.5 vs 61.3.
  • GPT-5.6 Sol (max) is stronger in coding, composite, knowledge, long context, maths, reasoning, agents & tools, instruction following, multimodal.
  • Claude Haiku 5.5 (high) is 40× cheaper ($0.200 vs $8.00 per 1M blended).
  • Claude Haiku 5.5 (high) streams 2.4× faster (174 vs 74 tokens per second).
MetricClaude Haiku 5.5 (high)GPT-5.6 Sol (max)
BenchLeader Index61.367.5
Coding score62.665.8
Composite score72.475.4
Knowledge score64.365.8
Long context score63.867.3
Maths score66.271.5
Reasoning score72.172.7
Agents & tools score–66.5
Instruction following score–71.0
Multimodal score–66.7
Blended price $/M$0.200$8.00
Output speed174 tok/s74 tok/s
Time to first answer22.6 s85.3 s
Context window1M1.1M
GPQA Diamond–93.5%
FrontierMath Tiers 1–3–89.1%
FrontierMath Tier 4–82.9%
OTIS Mock AIME97.2%100.0%
SimpleQA Verified–69.7%
Terminal-Bench–37.3%
OSWorld-Verified 2.0–27.3%
SciCode–57.1%
WeirdML–87.0%
ProofBench–83.0%
LMArena WebDev1587–
LiveBench–81.0%
LiveBench Reasoning–91.7%
LiveBench Coding–83.9%
LiveBench Agentic Coding–56.2%
LiveBench Mathematics–96.2%
LiveBench Data Analysis–79.8%
LiveBench Language–87.7%
LiveBench Instruction Following–71.8%
AA Intelligence Index v4.3.237.847.0
IFBench–72.7%
AA-LCR77.3%84.0%
MMMU-Pro–83.4%
AA-Omniscience5.822.0
Terminal-Bench Hard–65.9%
GPQA Diamond (AA)–94.1%
Humanity's Last Exam (AA)37.3%49.5%
SciCode (AA)48.7%57.1%
τ²-Bench Telecom (AA)–85.1%
LiveCodeBench–82.6%
MMLU-Pro–89.1%
IOI–91.2%
LegalBench–87.0%
CorpFin–64.4%
TaxEval–74.8%
Terminal-Bench 2.1 (Vals)–85.8%
SWE-bench (Vals)–96.2%
GPQA Diamond (Vals)–95.2%
Vals Index–58.0
PRBench Finance–50.5%
PRBench Legal–50.5%
ARC-AGI-1–96.5%
ARC-AGI-2–92.5%
ARC-AGI-3–7.8%
CritPt18.6%32.3%
GDPval-AA v2.145.9%55.6%
τ³-Banking (AA)–44.3%
ITBench SRE (AA)–56.2%
Analyst Agent (AA)–47.5%
BioMysteryBench–71.1%
Code Migration–52.9%
CUA-bench–8.3%
CyberBench–76.3%
Excel Modeling Benchmark–72.3%
Finance Agent v2–53.8%
Harvey's Legal Agent Benchmark–2.5%
Legal Research Bench–48.1%
MedCode–44.0%
MedScribe–85.2%
MMMU-Pro (Vals)–88.8%
MortgageTax–67.3%
MysteryMechanism–33.3%
ProgramBench–1.5%
Public Benefits Bench–66.5%
SAGE–52.6%
SkillsBench–54.1%
SREBench–30.5%
Tax Agent Bench–68.0%
Terminal-Bench 4.0 (Vals)–37.9%
Terminal-Bench Science–20.0%
Time Horizon Index: KSP–23.8%
Vals Multimodal Index–72.6%
Vibe Code Bench 1-100–20.0%
Vibe Code Bench v1.1–80.5%
Web Search Index–43.6%
Chess Puzzles–55.0%
EBR-bench–44.8%
Mystery Game Puzzles–58.0%
BALROG–60.0%
PostTrainBench–36.2%
DeepSWE v1.1–72.7%
LMCA–58.4%
DTBench–95.5%
CursorBench–41.7%
ALE-Bench–2176.9
GDP.pdf–30.7%
FrontierSWE–32.2%
BTF-3–13.7%
Terminal-Bench 4.0 (AA)21.7%39.9%
Terminal-Bench 2.1 (AA)–88.0%
AutomationBench33.7%–
GDP.pdf17.2%–
Harvey LAB1.1%–
AA-Omniscience: accuracy34.8%59.4%
AA-Omniscience: non-hallucination55.5%7.8%
AA-Briefcase v1.11442–

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Claude Haiku 5.5 vs GPT-5.6 Sol: questions

Is Claude Haiku 5.5 better than GPT-5.6 Sol?
GPT-5.6 Sol (max) leads on quality: 67.5 vs 61.3. The BenchLeader Index combines every independent quality benchmark; GPT-5.6 Sol (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Claude Haiku 5.5 better than GPT-5.6 Sol for coding?
GPT-5.6 Sol scores higher in coding (66 vs 63 on the category index, where 50 is average).
Which is cheaper, Claude Haiku 5.5 or GPT-5.6 Sol?
Claude Haiku 5.5 is cheaper: $0.200 against $8.00 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Haiku 5.5 or GPT-5.6 Sol?
Claude Haiku 5.5 streams faster: 174 against 74 output tokens per second.
Which has the larger context window?
GPT-5.6 Sol accepts more context: 1.1M against 1M tokens.