Claude 3.5 Haiku
Best configuration ranks #518 of 610 on the BenchLeader Index at 39.8 ±3.7. Last measured 2 Sept 2026. Released 4 Nov 2024.
- Blended price
- –
- Output speed
- –
- First answer
- –
- Context
- 200k
- Overall index40
- Reasoning37
- Coding35
- Agents & tools36
- Maths33
- Knowledge37
- Instruction following45
- Human preference49
- Multimodal34
- Long context39
- Composite40
Benchmark results
One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.
Reasoning
| Benchmark | default | Source |
|---|---|---|
| GPQA Diamond | 38.1%#233 | Epoch AI Benchmarking Hub |
| LMArena Hard Prompts | 1344#197 | LMArena |
| GPQA Diamond (AA)not in index | 40.8%#451 | Artificial Analysis |
| Humanity's Last Exam (AA)not in index | 3.6%#495 | Artificial Analysis |
| GPQA Diamond (Vals)not in index | 37.9%#123 | Vals AI |
Coding
| Benchmark | default | Source |
|---|---|---|
| SciCode | 27.4%#145 | SciCode |
| WeirdML | 30.7%#126 | WeirdML |
| LMArena Coding | 1386#187 | LMArena |
| LiveCodeBench | 41.9%#122 | Vals AI |
| Aider Polyglot | 28.0%#33 | Aider polyglot leaderboard |
Agents & tools
| Benchmark | default | Source |
|---|---|---|
| Terminal-Bench Hard | 2.3%#304 | Artificial Analysis |
| τ²-Bench Telecom (AA)not in index | 24.6%#289 | Artificial Analysis |
Maths
| Benchmark | default | Source |
|---|---|---|
| OTIS Mock AIME | 4.3%#229 | Epoch AI Benchmarking Hub |
| MATH Level 5 | 46.4%#56 | Epoch AI Benchmarking Hub |
| AIME (Vals) | 3.3%#89 | Vals AI |
| MGSM | 84.6%#65 | Vals AI |
Knowledge
| Benchmark | default | Source |
|---|---|---|
| AA-Omniscience | -22.5#192 | Artificial Analysis |
| MMLU-Pro | 64.1%#123 | Vals AI |
| LegalBench | 70.3%#113 | Vals AI |
| CorpFin | 50.8%#101 | Vals AI |
| TaxEval | 57.4%#130 | Vals AI |
Instruction following
| Benchmark | default | Source |
|---|---|---|
| IFBench | 42.8%#229 | Artificial Analysis |
Human preference
| Benchmark | default | Source |
|---|---|---|
| LMArena Text | 1324#207 | LMArena |
Multimodal
| Benchmark | default | Source |
|---|---|---|
| LMArena Vision | 1092#108 | LMArena |
| MMMU-Pro | 45.6%#218 | Artificial Analysis |
Long context
| Benchmark | default | Source |
|---|---|---|
| AA-LCR | 27.3%#333 | Artificial Analysis |
Composite
| Benchmark | default | Source |
|---|---|---|
| Epoch Capabilities Indexnot in index | 127.2#144 | Epoch AI Benchmarking Hub |
| AA Intelligence Index | 8.9#344 | Artificial Analysis |
See also
Data as of 9 Sept 2026. Compare with another model.