BenchLeader
AnthropicReasoning model

Claude Opus 4.5

Best configuration ranks #83 of 610 on the BenchLeader Index at 60.6 ±2.6 (thinking reasoning effort). Last measured 1 Sept 2026. Released 24 Nov 2025.

Blended price
$10.00/M
$5.00 in · $25.00 out
Output speed
46 tok/s
First answer
18 s
first token 1.34 s
Context
200k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index61
  2. Reasoning59
  3. Coding56
  4. Agents & tools74
  5. Maths64
  6. Knowledge61
  7. Instruction following55
  8. Multimodal58
  9. Long context65
  10. Composite66

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning45 tok/s1.34 s$0.0095Reasoning 39
medium45 tok/s1.34 s$0.0095Coding 65
high56.6#15445 tok/s1.34 s$0.0095Agents & tools 54 · Coding 63 · Composite 44 · Human preference 67 · Reasoning 68
thinkingbest60.6#8346 tok/s18 s$0.0095Agents & tools 74 · Coding 56 · Composite 66 · Instruction following 55 · Knowledge 61 · Long context 65 · Maths 64 · Multimodal 58 · Reasoning 59
default58.0#13245 tok/s1.34 s$0.0095Agents & tools 68 · Coding 57 · Composite 59 · Human preference 67 · Instruction following 45 · Knowledge 57 · Long context 62 · Maths 51 · Multimodal 58 · Reasoning 62

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoningmediumhighthinkingdefaultSourceTrend
GPQA Diamond86.0%#71Epoch AI Benchmarking Hub
Humanity's Last Exam25.2%#13Scale AI / CAIS
SimpleBench62.0%#22SimpleBench
LMArena Hard Prompts1499#251498#28LMArena
LiveBench Reasoningnot in index80.1%#38LiveBench
GPQA Diamond (AA)not in index86.6%#10981.0%#180Artificial Analysis
Humanity's Last Exam (AA)not in index30.1%#10213.2%#210Artificial Analysis
GPQA Diamond (Vals)not in index85.9%#4379.5%#68Vals AI
Kagi LLM Benchmark80.2%#870.7%#25Kagi LLM Benchmark
ARC-AGI-140.0%#13280.0%#81ARC Prize
ARC-AGI-27.8%#11037.6%#81ARC Prize

Coding

Benchmarkno reasoningmediumhighthinkingdefaultSourceTrend
SWE-bench Verified (Epoch)76.7%#9Epoch AI Benchmarking Hub
WeirdML63.7%#38WeirdML
GSO-Bench26.5%#10GSO-Bench
LMArena Coding1531#131523#21LMArena
LMArena WebDev1495#411468#48LMArena
LiveBench Codingnot in index79.7%#14LiveBench
LiveCodeBench83.7%#4575.0%#78Vals AI
IOI20.3%#2323.6%#19Vals AI
SWE-bench (Vals)not in index76.4%#35Vals AI
SWE-Bench Pro45.9%#6Scale AI SEAL
SWE-bench Verified (bash only)74.4%#576.8%#1SWE-bench
SWE-bench Verified (any scaffold)not in index79.2%#1SWE-bench

Agents & tools

Benchmarkno reasoningmediumhighthinkingdefaultSourceTrend
Terminal-Bench63.1%#12Terminal-Bench
GDPval45.5%#2OpenAI
Cybench82.0%#2Cybench
Remote Labor Index3.8%#6Scale AI / CAIS
APEX-Agents20.7%#35Mercor
LiveBench Agentic Codingnot in index39.7%#49LiveBench
Terminal-Bench Hard47.0%#2740.9%#51Artificial Analysis
τ²-Bench Telecom (AA)not in index89.5%#6886.3%#87Artificial Analysis
MCP Atlas69.8%#19Scale AI SEAL
BFCL Overall77.5%#1Berkeley Function Calling Leaderboard
τ²-bench85.3%#2τ²-bench

Maths

Benchmarkno reasoningmediumhighthinkingdefaultSourceTrend
FrontierMath Tiers 1–334.4%#53Epoch AI Benchmarking Hub
FrontierMath Tier 44.9%#51Epoch AI Benchmarking Hub
OTIS Mock AIME86.1%#79Epoch AI Benchmarking Hub
ProofBench36.0%#25Vals AI
LiveBench Mathematicsnot in index90.4%#23LiveBench
AIME (Vals)95.4%#1276.9%#50Vals AI
MGSM95.2%#194.8%#2Vals AI

Knowledge

Benchmarkno reasoningmediumhighthinkingdefaultSourceTrend
SimpleQA Verified45.7%#31Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index74.4%#28LiveBench
AA-Omniscience14#56-4.1#112Artificial Analysis
MMLU-Pro87.3%#3185.6%#55Vals AI
LegalBench84.6%#2882.8%#57Vals AI
CorpFin65.1%#3461.3%#54Vals AI
TaxEval74.9%#2574.3%#37Vals AI
MedQA95.9%#1093.2%#24Vals AI
PRBench Finance46.2%#16Scale AI SEAL
PRBench Legal44.2%#16Scale AI SEAL
MultiNRC48.6%#1241.2%#18Scale AI SEAL

Instruction following

Benchmarkno reasoningmediumhighthinkingdefaultSourceTrend
LiveBench Languagenot in index81.3%#20LiveBench
IFBench58.0%#13143.0%#225Artificial Analysis
MultiChallenge59.0%#12Scale AI SEAL
TutorBench51.2%#1649.8%#18Scale AI SEAL

Human preference

Benchmarkno reasoningmediumhighthinkingdefaultSourceTrend
LMArena Text1473#341469#40LMArena

Multimodal

Benchmarkno reasoningmediumhighthinkingdefaultSourceTrend
MMMU-Pro74.0%#8871.2%#108Artificial Analysis
VISTA46.4%#2145.3%#27Scale AI SEAL
MMMU (validation)80.7%#5MMMU
MMMU-Pro (official)not in index73.9%#8MMMU

Long context

Benchmarkno reasoningmediumhighthinkingdefaultSourceTrend
AA-LCR77.3%#9170.7%#154Artificial Analysis

Composite

Benchmarkno reasoningmediumhighthinkingdefaultSourceTrend
Epoch Capabilities Indexnot in index150.1#43Epoch AI Benchmarking Hub
LiveBench72.6%#39LiveBench
AA Intelligence Index29.1#8523.7#128Artificial Analysis

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Azure71 tok/s1.28 s$5.00$25.00200k
Google Vertex56 tok/s1.26 s$5.00$25.00200k
Claude Platform on AWS53 tok/s1.67 s$5.00$25.00200k
Amazon Bedrock36 tok/s1.75 s$5.00$25.00200k
Anthropic13 tok/s1.22 s$5.00$25.00200k

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$7.56$15.13$22.69$30.25Jan 26Feb 26Mar 26Apr 26May 26Jun 26Jul 26Aug 26
input outputnow $5.50 in · $27.50 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.500 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0095$0.008124.1 s
Summarise a 30-page report12,000 / 600$0.075$0.03530.6 s
Code edit6,000 / 1,500$0.068$0.04750.4 s
Agentic coding session60,000 / 4,000$0.400$0.1981.8 min
Structured extraction2,000 / 200$0.015$0.008321.9 s

See also

Data as of 9 Sept 2026. Compare these configurations.