BenchLeader
AnthropicReasoning modelAuto-detected

Claude Opus 4.7

Best configuration ranks #25 of 610 on the BenchLeader Index at 65.5 ±2.5. Last measured 8 Sept 2026. Released 16 Apr 2026.

Blended price
$10.00/M
$5.00 in · $25.00 out
Output speed
46 tok/s
First answer
21 s
first token 1.55 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index66
  2. Reasoning66
  3. Coding65
  4. Agents & tools69
  5. Maths65
  6. Knowledge66
  7. Instruction following59
  8. Human preference69
  9. Multimodal66
  10. Long context66
  11. Composite80

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning60.1#9444 tok/s1.57 s$0.0095Agents & tools 81 · Composite 68 · Instruction following 46 · Knowledge 71 · Long context 64 · Multimodal 61
high61.3#7146 tok/s21 s$0.0095Coding 68 · Human preference 71 · Multimodal 68 · Reasoning 55
xhigh57.2#14646 tok/s21 s$0.0095Composite 56 · Knowledge 60 · Maths 60 · Reasoning 65
max59.3#10746 tok/s21 s$0.0095Agents & tools 57 · Coding 67 · Maths 61 · Reasoning 63
thinking46 tok/s21 s$0.0095Reasoning 73
defaultbest65.5#2546 tok/s21 s$0.0095Agents & tools 69 · Coding 65 · Composite 80 · Human preference 69 · Instruction following 59 · Knowledge 66 · Long context 66 · Maths 65 · Multimodal 66 · Reasoning 66

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninghighxhighmaxthinkingdefaultSourceTrend
GPQA Diamond90.2%#3786.4%#66Epoch AI Benchmarking Hub
Humanity's Last Exam36.2%#7Scale AI / CAIS
SimpleBench61.7%#23SimpleBench
LMArena Hard Prompts1526#41519#6LMArena
LiveBench Reasoningnot in index87.2%#21LiveBench
GPQA Diamond (AA)not in index88.5%#8891.4%#47Artificial Analysis
Humanity's Last Exam (AA)not in index33.3%#9242.3%#41Artificial Analysis
GPQA Diamond (Vals)not in index90.2%#25Vals AI
Kagi LLM Benchmark80.7%#773.3%#18Kagi LLM Benchmark
ARC-AGI-30.2%#30ARC Prize

Coding

Benchmarkno reasoninghighxhighmaxthinkingdefaultSourceTrend
SWE-bench Verified (Epoch)83.5%#1Epoch AI Benchmarking Hub
SciCode54.5%#29SciCode
WeirdML76.4%#2275.5%#2576.4%#23WeirdML
FrontierCode38.5%#16Cognition
GSO-Bench44.1%#244.1%#2GSO-Bench
LMArena Coding1552#11547#4LMArena
LMArena WebDev1555#241557#22LMArena
LiveBench Codingnot in index82.1%#7LiveBench
LiveCodeBench85.1%#34Vals AI
IOI47.1%#10Vals AI
SWE-bench (Vals)not in index82.0%#21Vals AI

Agents & tools

Benchmarkno reasoninghighxhighmaxthinkingdefaultSourceTrend
Terminal-Bench80.2%#3Terminal-Bench
OSWorld-Verified 2.018.2%#8OSWorld
APEX-Agents33.9%#21Mercor
LiveBench Agentic Codingnot in index50.7%#29LiveBench
Terminal-Bench Hard54.5%#1551.5%#21Artificial Analysis
τ²-Bench Telecom (AA)not in index74.0%#13988.6%#73Artificial Analysis
Terminal-Bench 2.1 (Vals)68.5%#24Vals AI
MCP Atlas79.1%#11Scale AI SEAL
HiL-Bench41.7%#5Scale AI SEAL

Maths

Benchmarkno reasoninghighxhighmaxthinkingdefaultSourceTrend
FrontierMath Tiers 1–370.2%#17Epoch AI Benchmarking Hub
FrontierMath Tier 431.7%#25Epoch AI Benchmarking Hub
OTIS Mock AIME97.8%#1686.7%#71Epoch AI Benchmarking Hub
ProofBench54.0%#15Vals AI
LiveBench Mathematicsnot in index92.8%#15LiveBench
AIME (Vals)96.3%#7Vals AI
AIME 202695.8%#13MathArena
HMMT February 202693.9%#10MathArena
MathArena Apex40.6%#7MathArena

Knowledge

Benchmarkno reasoninghighxhighmaxthinkingdefaultSourceTrend
SimpleQA Verified51.7%#19Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index78.3%#17LiveBench
AA-Omniscience14.8#5327.3#25Artificial Analysis
MMLU-Pro89.9%#8Vals AI
LegalBench85.3%#21Vals AI
CorpFin66.1%#22Vals AI
TaxEval75.3%#20Vals AI

Instruction following

Benchmarkno reasoninghighxhighmaxthinkingdefaultSourceTrend
LiveBench Languagenot in index77.9%#30LiveBench
IFBench43.6%#21758.6%#128Artificial Analysis

Human preference

Benchmarkno reasoninghighxhighmaxthinkingdefaultSourceTrend
LMArena Text1502#41495#7LMArena
EQ-Bench 41311#5EQ-Bench

Multimodal

Benchmarkno reasoninghighxhighmaxthinkingdefaultSourceTrend
LMArena Vision1316#41317#3LMArena
MMMU-Pro76.4%#6378.8%#47Artificial Analysis

Long context

Benchmarkno reasoninghighxhighmaxthinkingdefaultSourceTrend
AA-LCR75.7%#10778.7%#75Artificial Analysis

Composite

Benchmarkno reasoninghighxhighmaxthinkingdefaultSourceTrend
Epoch Capabilities Indexnot in index156.3#16Epoch AI Benchmarking Hub
LiveBench76.5%#20LiveBench
AA Intelligence Index30.9#7440.7#32Artificial Analysis
Vals Indexnot in index56.1#19Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Google Vertex69 tok/s1.50 s$5.00$25.001M
Amazon Bedrock (Global)52 tok/s1.60 s$5.00$25.001M
Anthropic50 tok/s1.70 s$5.00$25.001M
Claude Platform on AWS43 tok/s1.25 s$5.00$25.001M

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$7.56$15.13$22.69$30.25Apr 26May 26Jun 26Jul 26Aug 26
input outputnow $5.50 in · $27.50 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.500 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0095$0.008127.7 s
Summarise a 30-page report12,000 / 600$0.075$0.03534.2 s
Code edit6,000 / 1,500$0.068$0.04753.7 s
Agentic coding session60,000 / 4,000$0.400$0.1981.8 min
Structured extraction2,000 / 200$0.015$0.008325.5 s

See also

Data as of 9 Sept 2026. Compare these configurations.