BenchLeader

Claude Opus 5 vs Claude Sonnet 5.5

Verdict
  • Claude Opus 5 (high) and Claude Sonnet 5.5 (max) are level on quality (68.3 vs 67.6).
  • Claude Opus 5 (high) is stronger in composite, human preference, knowledge, multimodal.
  • Claude Sonnet 5.5 (max) is stronger in agents & tools, coding, long context, maths, reasoning.
  • Claude Sonnet 5.5 (max) is 2.5× cheaper ($4.00 vs $10.00 per 1M blended).
  • Claude Sonnet 5.5 (max) streams 2.6× faster (141 vs 54 tokens per second).
MetricClaude Opus 5 (high)Claude Sonnet 5.5 (max)
BenchLeader Index68.367.6
Agents & tools score65.169.5
Coding score69.269.8
Composite score84.172.6
Human preference score68.9–
Knowledge score76.965.4
Long context score64.766.6
Maths score71.172.9
Multimodal score66.4–
Reasoning score72.781.1
Blended price $/M$10.00$4.00
Output speed54 tok/s141 tok/s
Time to first answer20.4 s487.1 s
Context window1M1M
GPQA Diamond–95.6%
FrontierMath Tiers 1–3–88.8%
FrontierMath Tier 4–80.5%
OTIS Mock AIME–100.0%
SimpleQA Verified–46.5%
Terminal-Bench50.3%61.8%
OSWorld-Verified 2.029.0%–
SciCode54.3%61.0%
WeirdML91.6%–
APEX-Agents–75.5%
ProofBench–100.0%
LMArena Text1490–
LMArena Hard Prompts1515–
LMArena Coding1534–
LMArena WebDev1657–
LMArena Vision1319–
LMArena Agent812
LiveBench–75.7%
LiveBench Reasoning–91.6%
LiveBench Coding–91.4%
LiveBench Agentic Coding–56.3%
LiveBench Mathematics–96.1%
LiveBench Data Analysis–59.5%
LiveBench Language–78.0%
LiveBench Instruction Following–56.8%
AA Intelligence Index v4.3.248.156
AA-LCR79.0%82.7%
MMMU-Pro82.4%–
AA-Omniscience33.732.3
GPQA Diamond (AA)93.7%–
Humanity's Last Exam (AA)52.8%55.0%
SciCode (AA)55.4%61.0%
ARC-AGI-197.5%–
ARC-AGI-288.3%–
ARC-AGI-330.2%–
CritPt28.3%31.4%
GDPval-AA v2.154.7%67.0%
τ³-Banking (AA)44.7%–
Analyst Agent (AA)–57.5%
LMArena Maths1521–
LMArena Creative Writing1471–
LMArena Instruction Following1497–
LMArena Multi-turn1484–
LMArena Longer Queries1504–
LMArena Document1490–
Mystery Game Puzzles–65.0%
DeepSWE v1.172.8%–
LMCA63.7%–
DTBench97.9%–
CursorBench44.7%55.5%
ALE-Bench2164.6–
FrontierSWE–61.9%
Terminal-Bench 4.0 (AA)46.0%63.6%
Terminal-Bench 2.1 (AA)87.6%–
AutomationBench53.6%71.8%
GDP.pdf19.6%25.8%
MLCR59.4%75.0%
Harvey LAB–2.8%
AA-Omniscience: accuracy58.9%54.0%
AA-Omniscience: non-hallucination38.8%53.0%
AA-Briefcase v1.115611823

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Claude Opus 5 vs Claude Sonnet 5.5: questions

Is Claude Opus 5 better than Claude Sonnet 5.5?
Claude Opus 5 (high) and Claude Sonnet 5.5 (max) are level on quality (68.3 vs 67.6). The BenchLeader Index combines every independent quality benchmark; Claude Opus 5 (high) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Claude Opus 5 better than Claude Sonnet 5.5 for coding?
Claude Sonnet 5.5 scores higher in coding (70 vs 69 on the category index, where 50 is average).
Is Claude Opus 5 better than Claude Sonnet 5.5 for agentic tasks?
Claude Sonnet 5.5 scores higher in agentic tasks (70 vs 65 on the category index, where 50 is average).
Which is cheaper, Claude Opus 5 or Claude Sonnet 5.5?
Claude Sonnet 5.5 is cheaper: $4.00 against $10.00 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Opus 5 or Claude Sonnet 5.5?
Claude Sonnet 5.5 streams faster: 141 against 54 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.