BenchLeader

Claude Sonnet 5 vs o3

Verdict
  • Claude Sonnet 5 (high) and o3 are level on quality (61.3 vs 61.6).
  • Claude Sonnet 5 (high) is stronger in composite, human preference, knowledge, long context, multimodal, reasoning.
  • o3 is stronger in agents & tools, coding, instruction following, maths.
  • They cost about the same ($3.50 per 1M blended).
  • o3 streams 1.8× faster (107 vs 61 tokens per second).
MetricClaude Sonnet 5 (high)o3
BenchLeader Index61.361.6
Agents & tools score60.870.3
Coding score61.562.0
Composite score69.554.5
Human preference score65.962.3
Knowledge score62.558.0
Long context score64.863.8
Multimodal score63.157.9
Reasoning score66.561.5
Instruction following score65.0
Maths score78.6
Blended price $/M$4.00$3.50
Output speed61 tok/s107 tok/s
Time to first answer2.0 s6.3 s
Context window1M200k
SciCode48.6%
WeirdML68.8%
Epoch Capabilities Index146.9
LMArena Text14611432
LMArena Hard Prompts14881441
LMArena Coding15211460
LMArena WebDev1537
LMArena Vision12781214
LMArena Agent5.9
AA Intelligence Index32.020.2
IFBench71.4%
AA-LCR76.7%74.7%
MMMU-Pro70.1%
AA-Omniscience-3.7-15.6
Terminal-Bench Hard37.1%
GPQA Diamond (AA)82.7%
Humanity's Last Exam (AA)35.7%20.1%
SciCode (AA)54.3%
τ²-Bench Telecom (AA)80.7%
PRBench Finance47.7%
PRBench Legal48.6%
MMMU (validation)82.9%
MMMU-Pro (official)76.4%
Kagi LLM Benchmark67.6%
IFEval (HELM)86.9%
Omni-MATH (HELM)71.4%
WildBench (HELM)86.1%
MMLU-Pro (HELM)85.9%
GPQA Diamond (HELM)75.3%
HELM Capabilities mean81.1%
Aider Polyglot81.3%
SWE-bench Verified (bash only)58.4%
SWE-bench Verified (any scaffold)58.4%
BFCL Overall63.0%

Data as of 2026-09-12. Best configuration of each model; every score links to its source on the model pages.

Claude Sonnet 5 vs o3: questions

Is Claude Sonnet 5 better than o3?
Claude Sonnet 5 (high) and o3 are level on quality (61.3 vs 61.6). The BenchLeader Index combines every independent quality benchmark; o3 is ahead overall as of 2026-09-12, but check the category scores for your use.
Is Claude Sonnet 5 better than o3 for coding?
o3 scores higher in coding (62 vs 62 on the category index, where 50 is average).
Is Claude Sonnet 5 better than o3 for agentic tasks?
o3 scores higher in agentic tasks (70 vs 61 on the category index, where 50 is average).
Which is cheaper, Claude Sonnet 5 or o3?
o3 is cheaper: $3.50 against $4.00 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Sonnet 5 or o3?
o3 streams faster: 107 against 61 output tokens per second.
Which has the larger context window?
Claude Sonnet 5 accepts more context: 1M against 200k tokens.