BenchLeader
AnthropicReasoning model

Claude Opus 5

Best configuration ranks #5 of 610 on the BenchLeader Index at 70.0 ±4.3. Last measured 5 Sept 2026. Released 24 Jul 2026.

Blended price
$10.00/M
$5.00 in · $25.00 out
Output speed
52 tok/s
First answer
82 s
first token 3.73 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index70
  2. Reasoning71
  3. Coding74
  4. Agents & tools72
  5. Maths67
  6. Knowledge71
  7. Human preference78
  8. Multimodal69
  9. Long context66
  10. Composite93

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
low61.7#6250 tok/s2.78 s$0.0095Agents & tools 57 · Coding 55 · Composite 79 · Knowledge 78 · Long context 67 · Maths 65 · Multimodal 64 · Reasoning 64
medium63.1#5051 tok/s6.48 s$0.0095Agents & tools 62 · Coding 59 · Composite 86 · Knowledge 79 · Long context 68 · Multimodal 66
high69.3#1051 tok/s23 s$0.0095Agents & tools 70 · Coding 72 · Composite 90 · Human preference 70 · Knowledge 80 · Long context 66 · Multimodal 68 · Reasoning 68
xhigh66.3#2050 tok/s30 s$0.0095Agents & tools 68 · Coding 64 · Composite 92 · Knowledge 81 · Long context 67 · Multimodal 69
max66.8#1652 tok/s82 s$0.0095Agents & tools 65 · Coding 73 · Composite 67 · Human preference 69 · Knowledge 67 · Maths 74 · Reasoning 70
defaultbest70.0#552 tok/s82 s$0.0095Agents & tools 72 · Coding 74 · Composite 93 · Human preference 78 · Knowledge 71 · Long context 66 · Maths 67 · Multimodal 69 · Reasoning 71

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarklowmediumhighxhighmaxdefaultSourceTrend
GPQA Diamond87.9%#5393.9%#1192.9%#18Epoch AI Benchmarking Hub
SimpleBench80.6%#2SimpleBench
LMArena Hard Prompts1518#81514#11LMArena
LiveBench Reasoningnot in index91.2%#4LiveBench
GPQA Diamond (AA)not in index88.9%#8491.9%#4493.7%#1193.7%#1193.2%#22Artificial Analysis
Humanity's Last Exam (AA)not in index43.4%#3251.3%#1352.8%#1154.4%#854.9%#5Artificial Analysis
GPQA Diamond (Vals)not in index93.4%#7Vals AI
ARC-AGI-197.5%#697.5%#6ARC Prize
ARC-AGI-288.3%#1290.4%#6ARC Prize
ARC-AGI-330.2%#12ARC Prize

Coding

BenchmarklowmediumhighxhighmaxdefaultSourceTrend
SciCode48.0%#7250.7%#5454.3%#3155.0%#2755.7%#21SciCode
WeirdML91.6%#691.8%#586.3%#11WeirdML
FrontierCode53.4%#2Cognition
LMArena Coding1533#111525#19LMArena
LMArena WebDev1661#61688#3LMArena
LiveBench Codingnot in index81.5%#9LiveBench
SciCode (AA)not in index49.2%#7651.5%#5955.4%#3155.7%#2856.4%#22Artificial Analysis
LiveCodeBench89.0%#4Vals AI
IOI91.7%#1Vals AI
SWE-bench (Vals)not in index97.0%#1Vals AI

Agents & tools

BenchmarklowmediumhighxhighmaxdefaultSourceTrend
Terminal-Bench51.8%#24Terminal-Bench
OSWorld-Verified 2.022.3%#625.3%#529.0%#330.2%#231.4%#1OSWorld
APEX-Agents43.5%#5Mercor
LMArena Agent11.4#310.8#4LMArena
LiveBench Agentic Codingnot in index65.2%#2LiveBench
Terminal-Bench 2.1 (Vals)84.6%#4Vals AI
MCP Atlas85.8%#3Scale AI SEAL
HiL-Bench57.0%#2Scale AI SEAL

Maths

BenchmarklowmediumhighxhighmaxdefaultSourceTrend
FrontierMath Tiers 1–385.6%#7Epoch AI Benchmarking Hub
FrontierMath Tier 473.2%#13Epoch AI Benchmarking Hub
OTIS Mock AIME93.3%#4498.9%#1197.8%#17Epoch AI Benchmarking Hub
ProofBench99.0%#2Vals AI
LiveBench Mathematicsnot in index95.7%#7LiveBench

Knowledge

BenchmarklowmediumhighxhighmaxdefaultSourceTrend
SimpleQA Verified59.9%#13Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index74.5%#27LiveBench
AA-Omniscience28.6#2231.0#1633.7#1435.4#1237.1#11Artificial Analysis
MMLU-Pro91.6%#2Vals AI
LegalBench87.0%#7Vals AI
CorpFin73.2%#1Vals AI
TaxEval75.1%#22Vals AI

Instruction following

BenchmarklowmediumhighxhighmaxdefaultSourceTrend
LiveBench Languagenot in index88.7%#4LiveBench

Human preference

BenchmarklowmediumhighxhighmaxdefaultSourceTrend
LMArena Text1493#91488#14LMArena
EQ-Bench 41385#1EQ-Bench

Multimodal

BenchmarklowmediumhighxhighmaxdefaultSourceTrend
LMArena Vision1322#2LMArena
MMMU-Pro79.8%#3981.6%#2582.4%#1984.0%#1484.7%#8Artificial Analysis

Long context

BenchmarklowmediumhighxhighmaxdefaultSourceTrend
AA-LCR81.3%#3482.0%#2579.0%#7180.3%#4579.3%#65Artificial Analysis

Composite

BenchmarklowmediumhighxhighmaxdefaultSourceTrend
Epoch Capabilities Indexnot in index162.6#4Epoch AI Benchmarking Hub
LiveBench80.1%#7LiveBench
AA Intelligence Index39.8#3745.1#1948.2#1249.6#1050.7#7Artificial Analysis
Vals Indexnot in index67.2#2Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Anthropic Fast115 tok/s2.67 s$10.00$50.001M
Azure (US)69 tok/s3.85 s$5.00$25.001M
Amazon Bedrock64 tok/s3.60 s$5.00$25.001M
Amazon Bedrock (US)62 tok/s4.46 s$5.50$27.501M
Anthropic59 tok/s4.09 s$5.00$25.001M
Google Vertex57 tok/s3.14 s$5.00$25.001M
Claude Platform on AWS50 tok/s4.99 s$5.00$25.001M
Google Vertex (Europe)33 tok/s2.40 s$5.50$27.501M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.500 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0095$0.00811.5 min
Summarise a 30-page report12,000 / 600$0.075$0.0351.6 min
Code edit6,000 / 1,500$0.068$0.0471.8 min
Agentic coding session60,000 / 4,000$0.400$0.1982.6 min
Structured extraction2,000 / 200$0.015$0.00831.4 min

See also

Data as of 9 Sept 2026. Compare these configurations.