BenchLeader
AnthropicReasoning model

Claude Opus 4.8

Best configuration ranks #31 of 610 on the BenchLeader Index at 64.7 ±3.3. Last measured 8 Sept 2026. Released 28 May 2026.

Blended price
$10.00/M
$5.00 in · $25.00 out
Output speed
58 tok/s
First answer
29 s
first token 1.68 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index65
  2. Reasoning64
  3. Coding66
  4. Agents & tools65
  5. Knowledge66
  6. Instruction following62
  7. Human preference66
  8. Multimodal64
  9. Long context65
  10. Composite82

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning58 tok/s29 s$0.0095Coding 64 · Maths 61 · Reasoning 62
low58 tok/s29 s$0.0095Maths 67 · Reasoning 63
medium58 tok/s29 s$0.0095Coding 68 · Reasoning 66
high61.3#7058 tok/s29 s$0.0095Agents & tools 65 · Coding 65 · Human preference 68 · Multimodal 65 · Reasoning 62
xhigh58 tok/s29 s$0.0095Coding 73 · Reasoning 30
max59.4#10458 tok/s29 s$0.0095Agents & tools 54 · Coding 62 · Composite 55 · Knowledge 61 · Maths 70 · Reasoning 65
thinking58 tok/s29 s$0.0095Reasoning 80
defaultbest64.7#3158 tok/s29 s$0.0095Agents & tools 65 · Coding 66 · Composite 82 · Human preference 66 · Instruction following 62 · Knowledge 66 · Long context 65 · Multimodal 64 · Reasoning 64

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowmediumhighxhighmaxthinkingdefaultSourceTrend
GPQA Diamond85.3%#7688.4%#4991.0%#27Epoch AI Benchmarking Hub
SimpleBench64.8%#17SimpleBench
LMArena Hard Prompts1515#101503#21LMArena
LiveBench Reasoningnot in index89.2%#13LiveBench
GPQA Diamond (AA)not in index92.0%#41Artificial Analysis
Humanity's Last Exam (AA)not in index48.7%#18Artificial Analysis
GPQA Diamond (Vals)not in index92.4%#16Vals AI
EnigmaEval23.5%#5Scale AI SEAL
Kagi LLM Benchmark88.8%#2Kagi LLM Benchmark
ARC-AGI-188.0%#5791.5%#4492.0%#4092.5%#35ARC Prize
ARC-AGI-262.2%#5171.7%#3672.1%#34ARC Prize
ARC-AGI-31.5%#18ARC Prize

Coding

Benchmarkno reasoninglowmediumhighxhighmaxthinkingdefaultSourceTrend
SciCode53.5%#41SciCode
WeirdML70.5%#2976.0%#2482.9%#14WeirdML
FrontierCode46.5%#7Cognition
GSO-Bench47.1%#1GSO-Bench
LMArena Coding1534#81527#17LMArena
LMArena WebDev1561#211540#27LMArena
LiveBench Codingnot in index81.8%#8LiveBench
SciCode (AA)not in index54.4%#38Artificial Analysis
LiveCodeBench87.8%#11Vals AI
SWE-bench (Vals)not in index88.6%#12Vals AI

Agents & tools

Benchmarkno reasoninglowmediumhighxhighmaxthinkingdefaultSourceTrend
Terminal-Bench23.6%#51Terminal-Bench
OSWorld-Verified 2.020.6%#7OSWorld
Remote Labor Index8.3%#2Scale AI / CAIS
APEX-Agents42.5%#6Mercor
LMArena Agent7.7#6LMArena
LiveBench Agentic Codingnot in index50.5%#30LiveBench
Terminal-Bench Hard58.3%#10Artificial Analysis
τ²-Bench Telecom (AA)not in index94.4%#31Artificial Analysis
Terminal-Bench 2.1 (Vals)71.9%#18Vals AI
MCP Atlas82.2%#7Scale AI SEAL
HiL-Bench35.3%#9Scale AI SEAL

Maths

Benchmarkno reasoninglowmediumhighxhighmaxthinkingdefaultSourceTrend
FrontierMath Tiers 1–380.0%#11Epoch AI Benchmarking Hub
FrontierMath Tier 456.1%#18Epoch AI Benchmarking Hub
OTIS Mock AIME84.4%#8397.8%#1798.3%#14Epoch AI Benchmarking Hub
ProofBench69.0%#9Vals AI
LiveBench Mathematicsnot in index94.3%#10LiveBench
AIME 2026100.0%#1MathArena
HMMT February 202695.5%#6MathArena
MathArena Apex81.3%#1MathArena

Knowledge

Benchmarkno reasoninglowmediumhighxhighmaxthinkingdefaultSourceTrend
SimpleQA Verified53.0%#16Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index66.0%#44LiveBench
AA-Omniscience28.8#20Artificial Analysis
MMLU-Pro89.6%#9Vals AI
LegalBench83.6%#48Vals AI
CorpFin66.7%#17Vals AI
TaxEval75.6%#14Vals AI

Instruction following

Benchmarkno reasoninglowmediumhighxhighmaxthinkingdefaultSourceTrend
LiveBench Languagenot in index79.7%#26LiveBench
IFBench62.2%#116Artificial Analysis

Human preference

Benchmarkno reasoninglowmediumhighxhighmaxthinkingdefaultSourceTrend
LMArena Text1482#181473#33LMArena
EQ-Bench 41281#6EQ-Bench

Multimodal

Benchmarkno reasoninglowmediumhighxhighmaxthinkingdefaultSourceTrend
LMArena Vision1294#191289#23LMArena

Long context

Benchmarkno reasoninglowmediumhighxhighmaxthinkingdefaultSourceTrend
AA-LCR77.7%#87Artificial Analysis

Composite

Benchmarkno reasoninglowmediumhighxhighmaxthinkingdefaultSourceTrend
Epoch Capabilities Indexnot in index158.3#10Epoch AI Benchmarking Hub
LiveBench76.2%#21LiveBench
AA Intelligence Index42.0#29Artificial Analysis
Vals Indexnot in index60.9#8Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Anthropic Fast126 tok/s1.63 s$10.00$50.001M
Amazon Bedrock (EU)78 tok/s1.31 s$5.50$27.501M
Amazon Bedrock68 tok/s2.38 s$5.00$25.001M
Claude Platform on AWS63 tok/s1.73 s$5.00$25.001M
Google Vertex63 tok/s2.78 s$5.00$25.001M
Anthropic62 tok/s1.43 s$5.00$25.001M

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$7.56$15.13$22.69$30.25May 26Jun 26Jul 26Aug 26
input outputnow $5.50 in · $27.50 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.500 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0095$0.008134.7 s
Summarise a 30-page report12,000 / 600$0.075$0.03539.9 s
Code edit6,000 / 1,500$0.068$0.04755.5 s
Agentic coding session60,000 / 4,000$0.400$0.1981.6 min
Structured extraction2,000 / 200$0.015$0.008333.0 s

See also

Data as of 9 Sept 2026. Compare these configurations.