BenchLeader

Claude Sonnet 5 vs Claude Sonnet 5.5

Verdict
  • Claude Sonnet 5.5 (max) leads on quality: 67.6 vs 60.5.
  • Claude Sonnet 5 (high) is stronger in human preference, multimodal.
  • Claude Sonnet 5.5 (max) is stronger in agents & tools, coding, composite, knowledge, long context, maths, reasoning.
  • They cost about the same ($4.00 per 1M blended).
  • Claude Sonnet 5.5 (max) streams 2.3× faster (141 vs 62 tokens per second).
MetricClaude Sonnet 5 (high)Claude Sonnet 5.5 (max)
BenchLeader Index60.567.6
Agents & tools score57.969.5
Coding score61.169.8
Composite score65.472.6
Human preference score65.3–
Knowledge score60.065.4
Long context score63.566.6
Maths score65.872.9
Multimodal score61.3–
Reasoning score66.381.1
Blended price $/M$4.00$4.00
Output speed62 tok/s141 tok/s
Time to first answer12.3 s487.1 s
Context window1M1M
GPQA Diamond–95.6%
FrontierMath Tiers 1–3–88.8%
FrontierMath Tier 4–80.5%
OTIS Mock AIME–100.0%
SimpleQA Verified–46.5%
Terminal-Bench–61.8%
SciCode54.3%61.0%
WeirdML68.8%–
APEX-Agents–75.5%
ProofBench–100.0%
LMArena Text1461–
LMArena Hard Prompts1490–
LMArena Coding1519–
LMArena WebDev1541–
LMArena Vision1274–
LMArena Agent4.912
LiveBench–75.7%
LiveBench Reasoning–91.6%
LiveBench Coding–91.4%
LiveBench Agentic Coding–56.3%
LiveBench Mathematics–96.1%
LiveBench Data Analysis–59.5%
LiveBench Language–78.0%
LiveBench Instruction Following–56.8%
AA Intelligence Index v4.3.231.756
AA-LCR76.7%82.7%
AA-Omniscience-3.732.3
Humanity's Last Exam (AA)35.7%55.0%
SciCode (AA)54.3%61.0%
CritPt15.1%31.4%
GDPval-AA v2.137.9%67.0%
Analyst Agent (AA)–57.5%
LMArena Maths1473–
LMArena Creative Writing1435–
LMArena Instruction Following1466–
LMArena Multi-turn1473–
LMArena Longer Queries1481–
LMArena Document1476–
Mystery Game Puzzles–65.0%
DeepSWE v1.148.2%–
LMCA50.0%–
DTBench84.5%–
CursorBench30.8%55.5%
ALE-Bench1463.1–
FrontierSWE–61.9%
Terminal-Bench 4.0 (AA)5.0%63.6%
AutomationBench–71.8%
GDP.pdf–25.8%
MLCR–75.0%
Harvey LAB–2.8%
AA-Omniscience: accuracy37.4%54.0%
AA-Omniscience: non-hallucination34.3%53.0%
AA-Briefcase v1.1–1823

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Claude Sonnet 5 vs Claude Sonnet 5.5: questions

Is Claude Sonnet 5 better than Claude Sonnet 5.5?
Claude Sonnet 5.5 (max) leads on quality: 67.6 vs 60.5. The BenchLeader Index combines every independent quality benchmark; Claude Sonnet 5.5 (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Claude Sonnet 5 better than Claude Sonnet 5.5 for coding?
Claude Sonnet 5.5 scores higher in coding (70 vs 61 on the category index, where 50 is average).
Is Claude Sonnet 5 better than Claude Sonnet 5.5 for agentic tasks?
Claude Sonnet 5.5 scores higher in agentic tasks (70 vs 58 on the category index, where 50 is average).
Which is cheaper, Claude Sonnet 5 or Claude Sonnet 5.5?
Claude Sonnet 5 is cheaper: $4.00 against $4.00 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Sonnet 5 or Claude Sonnet 5.5?
Claude Sonnet 5.5 streams faster: 141 against 62 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.