BenchLeader
AnthropicReasoning modelAuto-detected

Claude Opus 4.6

Best configuration ranks #48 of 610 on the BenchLeader Index at 63.1 ±3.6. Last measured 8 Sept 2026. Released 5 Feb 2026.

Blended price
$10.00/M
$5.00 in · $25.00 out
Output speed
38 tok/s
First answer
2.05 s
first token 2.13 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index63
  2. Reasoning62
  3. Coding62
  4. Agents & tools69
  5. Maths66
  6. Knowledge71
  7. Instruction following54
  8. Human preference65
  9. Multimodal64
  10. Long context66
  11. Composite69

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning55.3#17938 tok/s2.05 s$0.0095Instruction following 52 · Knowledge 66 · Multimodal 56
low38 tok/s2.05 s$0.0095Reasoning 63
medium38 tok/s2.05 s$0.0095Reasoning 64
high61.3#6938 tok/s2.05 s$0.0095Coding 67 · Composite 50 · Human preference 71 · Maths 61 · Multimodal 68 · Reasoning 67
max54.1#20638 tok/s2.05 s$0.0095Agents & tools 57 · Instruction following 33 · Knowledge 60 · Maths 59 · Multimodal 57 · Reasoning 63
thinking62.3#5938 tok/s2.05 s$0.0095Coding 63 · Knowledge 62 · Maths 65 · Reasoning 75
defaultbest63.1#4838 tok/s2.05 s$0.0095Agents & tools 69 · Coding 62 · Composite 69 · Human preference 65 · Instruction following 54 · Knowledge 71 · Long context 66 · Maths 66 · Multimodal 64 · Reasoning 62

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowmediumhighmaxthinkingdefaultSourceTrend
GPQA Diamond88.4%#4990.5%#35Epoch AI Benchmarking Hub
Humanity's Last Exam34.4%#819.0%#20Scale AI / CAIS
SimpleBench67.6%#16SimpleBench
LMArena Hard Prompts1533#11527#3LMArena
LiveBench Reasoningnot in index88.7%#15LiveBench
GPQA Diamond (AA)not in index89.6%#71Artificial Analysis
Humanity's Last Exam (AA)not in index39.9%#55Artificial Analysis
GPQA Diamond (Vals)not in index89.7%#28Vals AI
Kagi LLM Benchmark83.6%#572.4%#23Kagi LLM Benchmark
ARC-AGI-186.0%#7092.0%#4094.0%#2993.0%#33ARC Prize
ARC-AGI-264.6%#4966.3%#4669.2%#3868.8%#39ARC Prize
ARC-AGI-30.5%#22ARC Prize

Coding

Benchmarkno reasoninglowmediumhighmaxthinkingdefaultSourceTrend
SWE-bench Verified (Epoch)78.7%#4Epoch AI Benchmarking Hub
WeirdML78.0%#1877.9%#19WeirdML
FrontierCode26.6%#20Cognition
GSO-Bench41.2%#433.3%#7GSO-Bench
LMArena Coding1552#21546#5LMArena
LMArena WebDev1546#251537#29LMArena
LiveBench Codingnot in index78.2%#22LiveBench
LiveCodeBench84.7%#36Vals AI
SWE-bench (Vals)not in index78.2%#30Vals AI
SWE-Bench Pro51.9%#4Scale AI SEAL
SWE-bench Verified (bash only)75.6%#4SWE-bench
SWE-bench Verified (any scaffold)not in index75.6%#7SWE-bench

Agents & tools

Benchmarkno reasoninglowmediumhighmaxthinkingdefaultSourceTrend
Terminal-Bench79.8%#5Terminal-Bench
Cybench93.0%#1Cybench
Remote Labor Index4.2%#5Scale AI / CAIS
APEX-Agents32.1%#2532.4%#24Mercor
LiveBench Agentic Codingnot in index49.0%#33LiveBench
Terminal-Bench Hard48.5%#26Artificial Analysis
τ²-Bench Telecom (AA)not in index92.1%#53Artificial Analysis
MCP Atlas76.8%#14Scale AI SEAL
HiL-Bench38.3%#8Scale AI SEAL

Maths

Benchmarkno reasoninglowmediumhighmaxthinkingdefaultSourceTrend
FrontierMath Tiers 1–366.0%#21Epoch AI Benchmarking Hub
FrontierMath Tier 426.8%#31Epoch AI Benchmarking Hub
OTIS Mock AIME91.1%#5594.4%#39Epoch AI Benchmarking Hub
ProofBench50.0%#17Vals AI
LiveBench Mathematicsnot in index89.3%#27LiveBench
AIME (Vals)95.6%#8Vals AI
AIME 202696.7%#6MathArena
HMMT February 202696.2%#5MathArena
MathArena Apex34.5%#8MathArena

Knowledge

Benchmarkno reasoninglowmediumhighmaxthinkingdefaultSourceTrend
SimpleQA Verified47.0%#29Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index69.9%#41LiveBench
AA-Omniscience13.7#58Artificial Analysis
MMLU-Pro89.1%#15Vals AI
LegalBench85.3%#19Vals AI
CorpFin67.0%#13Vals AI
TaxEval76.0%#8Vals AI
MedQA95.4%#12Vals AI
PRBench Finance53.3%#3Scale AI SEAL
PRBench Legal52.3%#4Scale AI SEAL
MultiNRC48.3%#1357.1%#8Scale AI SEAL

Instruction following

Benchmarkno reasoninglowmediumhighmaxthinkingdefaultSourceTrend
LiveBench Languagenot in index83.3%#14LiveBench
IFBench53.1%#155Artificial Analysis
MultiChallenge56.0%#1737.1%#28Scale AI SEAL
TutorBench53.5%#1053.7%#8Scale AI SEAL

Human preference

Benchmarkno reasoninglowmediumhighmaxthinkingdefaultSourceTrend
LMArena Text1505#21498#6LMArena
EQ-Bench 41223#12EQ-Bench

Multimodal

Benchmarkno reasoninglowmediumhighmaxthinkingdefaultSourceTrend
LMArena Vision1315#51311#7LMArena
MMMU-Pro75.4%#71Artificial Analysis
VISTA45.5%#2546.1%#22Scale AI SEAL

Long context

Benchmarkno reasoninglowmediumhighmaxthinkingdefaultSourceTrend
AA-LCR78.0%#82Artificial Analysis

Composite

Benchmarkno reasoninglowmediumhighmaxthinkingdefaultSourceTrend
Epoch Capabilities Indexnot in index155.3#22Epoch AI Benchmarking Hub
LiveBench74.5%#31LiveBench
AA Intelligence Index31.9#71Artificial Analysis

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Amazon Bedrock58 tok/s1.92 s$5.00$25.001M
Google Vertex42 tok/s2.15 s$5.00$25.001M
Anthropic34 tok/s2.39 s$5.00$25.001M
Claude Platform on AWS32 tok/s1.78 s$5.00$25.001M
Azure29 tok/s2.13 s$5.00$25.001M

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$7.56$15.13$22.69$30.25Feb 26Mar 26Apr 26May 26Jun 26Jul 26Aug 26Sept 26
input outputnow $5.50 in · $27.50 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.500 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0095$0.008110.0 s
Summarise a 30-page report12,000 / 600$0.075$0.03517.9 s
Code edit6,000 / 1,500$0.068$0.04741.6 s
Agentic coding session60,000 / 4,000$0.400$0.1981.8 min
Structured extraction2,000 / 200$0.015$0.00837.3 s

See also

Data as of 9 Sept 2026. Compare these configurations.