BenchLeader

Claude Opus 4.8 vs Claude Opus 5

Verdict
  • Claude Opus 5 leads on quality: 70.0 vs 64.7.
  • Claude Opus 4.8 is stronger in instruction following.
  • Claude Opus 5 is stronger in agents & tools, coding, composite, human preference, knowledge, long context, multimodal, reasoning, maths.
  • They cost about the same ($10.00 per 1M blended).
MetricClaude Opus 4.8Claude Opus 5
BenchLeader Index64.770.0
Agents & tools score65.071.6
Coding score65.774.4
Composite score81.992.9
Human preference score65.678.2
Instruction following score61.6
Knowledge score66.370.5
Long context score65.366.2
Multimodal score64.469.4
Reasoning score64.271.1
Maths score67.1
Blended price $/M$10.00$10.00
Output speed58 tok/s52 tok/s
Time to first answer29.5 s81.6 s
Context window1M1M
GPQA Diamond92.9%
OTIS Mock AIME97.8%
SimpleBench64.8%80.6%
Remote Labor Index8.3%
WeirdML86.3%
FrontierCode46.5%
GSO-Bench47.1%
Epoch Capabilities Index158.3162.6
LMArena Text1473
LMArena Hard Prompts1503
LMArena Coding1527
LMArena WebDev1540
LMArena Vision1289
AA Intelligence Index42.050.7
IFBench62.2%
AA-LCR77.7%79.3%
MMMU-Pro84.7%
AA-Omniscience28.837.1
Terminal-Bench Hard58.3%
GPQA Diamond (AA)92.0%93.2%
Humanity's Last Exam (AA)48.7%54.9%
SciCode (AA)54.4%56.4%
τ²-Bench Telecom (AA)94.4%
LiveCodeBench87.8%89.0%
MMLU-Pro89.6%91.6%
IOI91.7%
LegalBench83.6%87.0%
CorpFin66.7%73.2%
TaxEval75.6%75.1%
Terminal-Bench 2.1 (Vals)71.9%84.6%
SWE-bench (Vals)88.6%97.0%
GPQA Diamond (Vals)92.4%93.4%
Vals Index60.967.2
HiL-Bench35.3%57.0%
EQ-Bench 412811385

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.