BenchLeader
AnthropicAuto-detected

Claude 3 Sonnet

Best configuration ranks #556 of 610 on the BenchLeader Index at 37.9 ±4.7. Last measured 2 Sept 2026. Released 29 Feb 2024.

Blended price
Output speed
First answer
Context
200k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index38
  2. Reasoning36
  3. Coding33
  4. Maths28
  5. Human preference44
  6. Multimodal28
  7. Composite36

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Coding

BenchmarkdefaultSource
WeirdML10.2%#144WeirdML
LMArena Coding1318#238LMArena

Human preference

BenchmarkdefaultSource
LMArena Text1281#247LMArena

Multimodal

BenchmarkdefaultSource
LMArena Vision984#128LMArena
MMMU (validation)53.1%#39MMMU

See also

Data as of 9 Sept 2026. Compare with another model.