BenchLeader
AnthropicReasoning model

Claude 3.7 Sonnet

Best configuration ranks #255 of 610 on the BenchLeader Index at 52.1 ±3.0 (thinking reasoning effort). Last measured 2 Sept 2026. Released 24 Feb 2025.

Blended price
$6.00/M
$3.00 in · $15.00 out
Output speed
First answer
Context
200k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index52
  2. Reasoning58
  3. Coding47
  4. Agents & tools52
  5. Maths49
  6. Knowledge55
  7. Instruction following46
  8. Human preference57
  9. Multimodal55
  10. Composite51

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning$0.0057Coding 59
thinkingbest52.1#255$0.0057Agents & tools 52 · Coding 47 · Composite 51 · Human preference 57 · Instruction following 46 · Knowledge 55 · Maths 49 · Multimodal 55 · Reasoning 58
default49.8#300$0.0057Agents & tools 48 · Coding 47 · Composite 48 · Human preference 55 · Instruction following 46 · Knowledge 54 · Long context 50 · Maths 51 · Multimodal 50 · Reasoning 49

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoningthinkingdefaultSource
GPQA Diamond79.7%#115Epoch AI Benchmarking Hub
Humanity's Last Exam8.0%#31Scale AI / CAIS
SimpleBench46.4%#51SimpleBench
LMArena Hard Prompts1417#1351398#154LMArena
GPQA Diamond (AA)not in index77.2%#21565.6%#326Artificial Analysis
Humanity's Last Exam (AA)not in index9.7%#2644.2%#452Artificial Analysis
GPQA Diamond (Vals)not in index75.3%#7867.4%#97Vals AI

Coding

Benchmarkno reasoningthinkingdefaultSource
SWE-bench Verified (Epoch)61.0%#29Epoch AI Benchmarking Hub
GSO-Bench3.8%#26GSO-Bench
LMArena Coding1452#1251431#151LMArena
LiveCodeBench60.4%#10256.7%#107Vals AI
Aider Polyglot60.4%#1464.9%#10Aider polyglot leaderboard
SWE-bench Verified (bash only)52.8%#32SWE-bench
SWE-bench Verified (any scaffold)not in index70.4%#16SWE-bench

Agents & tools

Benchmarkno reasoningthinkingdefaultSource
Cybench20.0%#11Cybench
Terminal-Bench Hard21.2%#15221.2%#152Artificial Analysis
τ²-Bench Telecom (AA)not in index54.7%#18050.0%#191Artificial Analysis

Maths

Benchmarkno reasoningthinkingdefaultSource
OTIS Mock AIME57.8%#150Epoch AI Benchmarking Hub
MATH Level 591.2%#18Epoch AI Benchmarking Hub
AIME (Vals)44.6%#6222.3%#77Vals AI
MGSM93.0%#1292.4%#19Vals AI

Knowledge

Benchmarkno reasoningthinkingdefaultSource
AA-Omniscience-0.7#99-9.7#136Artificial Analysis
MMLU-Pro82.7%#7480.7%#82Vals AI
LegalBench82.5%#6180.0%#83Vals AI
CorpFin60.4%#68Vals AI
TaxEval74.0%#4172.4%#62Vals AI
MedQA90.2%#46Vals AI
MultiNRC27.8%#26Scale AI SEAL

Instruction following

Benchmarkno reasoningthinkingdefaultSource
IFBench48.3%#18144.0%#214Artificial Analysis
MultiChallenge51.6%#21Scale AI SEAL

Human preference

Benchmarkno reasoningthinkingdefaultSource
LMArena Text1388#1441372#160LMArena

Multimodal

Benchmarkno reasoningthinkingdefaultSource
LMArena Vision1169#891151#98LMArena
MMMU-Pro60.1%#178Artificial Analysis
VISTA48.2%#1643.0%#35Scale AI SEAL
MMMU (validation)75.0%#14MMMU

Long context

Benchmarkno reasoningthinkingdefaultSource
Fiction.LiveBench 120k53.1%#20Fiction.live
AA-LCR51.7%#252Artificial Analysis

Composite

Benchmarkno reasoningthinkingdefaultSource
Epoch Capabilities Indexnot in index141.2#94Epoch AI Benchmarking Hub
AA Intelligence Index17.7#19715.3#219Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0057
Summarise a 30-page report12,000 / 600$0.045
Code edit6,000 / 1,500$0.041
Agentic coding session60,000 / 4,000$0.240
Structured extraction2,000 / 200$0.0090

See also

Data as of 9 Sept 2026. Compare these configurations.