BenchLeader
AnthropicReasoning model

Claude Haiku 4.5

Best configuration ranks #336 of 610 on the BenchLeader Index at 48.0 ±4.5 (thinking reasoning effort). Last measured 5 Sept 2026. Released 15 Oct 2025.

Blended price
$2.00/M
$1.00 in · $5.00 out
Output speed
83 tok/s
First answer
22 s
first token 0.77 s
Context
200k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index48
  2. Reasoning39
  3. Coding30
  4. Agents & tools45
  5. Maths58
  6. Knowledge52
  7. Instruction following48
  8. Multimodal43
  9. Long context64
  10. Composite51

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
high79 tok/s0.77 s$0.0019Coding 59
thinkingbest48.0#33683 tok/s22 s$0.0019Agents & tools 45 · Coding 30 · Composite 51 · Instruction following 48 · Knowledge 52 · Long context 64 · Maths 58 · Multimodal 43 · Reasoning 39
default47.9#34179 tok/s0.77 s$0.0019Agents & tools 45 · Coding 49 · Composite 48 · Human preference 50 · Instruction following 44 · Knowledge 46 · Long context 51 · Maths 61 · Multimodal 39 · Reasoning 44

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkhighthinkingdefaultSourceTrend
GPQA Diamond71.2%#154Epoch AI Benchmarking Hub
LMArena Hard Prompts1441#105LMArena
GPQA Diamond (AA)not in index67.2%#30764.7%#329Artificial Analysis
Humanity's Last Exam (AA)not in index10.4%#2514.2%#441Artificial Analysis
GPQA Diamond (Vals)not in index72.2%#85Vals AI
ARC-AGI-147.7%#12514.3%#163ARC Prize
ARC-AGI-24.0%#1271.3%#153ARC Prize

Coding

BenchmarkhighthinkingdefaultSourceTrend
SciCode43.3%#92SciCode
WeirdML45.4%#82WeirdML
LMArena Coding1481#90LMArena
LMArena WebDev1329#97LMArena
SciCode (AA)not in index42.3%#105Artificial Analysis
LiveCodeBench41.2%#124Vals AI
IOI6.2%#42Vals AI
SWE-bench (Vals)not in index66.6%#68Vals AI
SWE-Bench Pro39.5%#12Scale AI SEAL
SWE-bench Verified (bash only)66.6%#16SWE-bench
SWE-bench Verified (any scaffold)not in index66.6%#20SWE-bench

Agents & tools

BenchmarkhighthinkingdefaultSourceTrend
Terminal-Bench35.5%#41Terminal-Bench
APEX-Agents8.9%#53Mercor
Terminal-Bench Hard27.3%#12327.3%#123Artificial Analysis
τ²-Bench Telecom (AA)not in index54.7%#18032.5%#241Artificial Analysis
Terminal-Bench 2.1 (Vals)43.8%#53Vals AI
MCP Atlas40.2%#30Scale AI SEAL
BFCL Overall68.7%#6Berkeley Function Calling Leaderboard

Maths

BenchmarkhighthinkingdefaultSourceTrend
OTIS Mock AIME66.7%#134Epoch AI Benchmarking Hub
MATH Level 596.4%#11Epoch AI Benchmarking Hub
AIME (Vals)82.7%#44Vals AI
MGSM92.2%#21Vals AI

Knowledge

BenchmarkhighthinkingdefaultSourceTrend
SimpleQA Verified13.2%#68Epoch AI Benchmarking Hub
AA-Omniscience-4.4#114-7.6#126Artificial Analysis
MMLU-Pro78.7%#96Vals AI
LegalBench81.2%#75Vals AI
CorpFin60.6%#6660.3%#69Vals AI
TaxEval67.5%#103Vals AI
MedQA79.6%#71Vals AI

Instruction following

BenchmarkhighthinkingdefaultSourceTrend
IFBench54.3%#15142.0%#234Artificial Analysis
MultiChallenge50.5%#23Scale AI SEAL

Human preference

BenchmarkhighthinkingdefaultSourceTrend
LMArena Text1413#117LMArena
EQ-Bench 41064#25EQ-Bench

Multimodal

BenchmarkhighthinkingdefaultSourceTrend
MMMU-Pro58.5%#18355.1%#194Artificial Analysis

Long context

BenchmarkhighthinkingdefaultSourceTrend
AA-LCR74.3%#11649.7%#256Artificial Analysis

Composite

BenchmarkhighthinkingdefaultSourceTrend
Epoch Capabilities Indexnot in index142.4#88Epoch AI Benchmarking Hub
AA Intelligence Index17.6#19815.4#216Artificial Analysis
Vals Indexnot in index22.9#47Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Amazon Bedrock (EU)97 tok/s0.99 s$1.10$5.50200k
Google Vertex (Europe)74 tok/s0.47 s$1.10$5.50200k
Amazon Bedrock (US)72 tok/s1.36 s$1.10$5.50200k
Google Vertex67 tok/s0.71 s$1.00$5.00200k
Amazon Bedrock (Global)65 tok/s0.71 s$1.00$5.00200k
Anthropic59 tok/s0.70 s$1.00$5.00200k
Azure53 tok/s0.71 s$1.00$5.00200k

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$1.51$3.03$4.54$6.05Jan 26Feb 26Mar 26Apr 26May 26Jun 26Jul 26Aug 26
input outputnow $1.10 in · $5.50 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.100 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0019$0.001625.8 s
Summarise a 30-page report12,000 / 600$0.015$0.006929.4 s
Code edit6,000 / 1,500$0.013$0.009540.2 s
Agentic coding session60,000 / 4,000$0.080$0.0401.2 min
Structured extraction2,000 / 200$0.0030$0.001724.6 s

See also

Data as of 9 Sept 2026. Compare these configurations.