Llama 3.1 405B
Best configuration ranks #457 of 610 on the BenchLeader Index at 42.6 ±4.4. Last measured 2 Sept 2026. Released 23 Jul 2024.
- Blended price
- –
- Output speed
- –
- First answer
- –
- Context
- 128k
- Overall index43
- Reasoning37
- Coding34
- Agents & tools38
- Maths37
- Knowledge56
- Instruction following42
- Human preference50
- Long context38
- Composite38
Benchmark results
One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.
Reasoning
| Benchmark | default | Source |
|---|---|---|
| GPQA Diamond | 50.9%#198 | Epoch AI Benchmarking Hub |
| SimpleBench | 23.0%#74 | SimpleBench |
| LMArena Hard Prompts | 1341#198 | LMArena |
| GPQA Diamond (AA)not in index | 51.5%#399 | Artificial Analysis |
| Humanity's Last Exam (AA)not in index | 4.0%#469 | Artificial Analysis |
Coding
| Benchmark | default | Source |
|---|---|---|
| WeirdML | 21.4%#135 | WeirdML |
| LMArena Coding | 1376#194 | LMArena |
| SWE-Bench Pro | 11.2%#21 | Scale AI SEAL |
Agents & tools
| Benchmark | default | Source |
|---|---|---|
| Cybench | 7.5%#18 | Cybench |
| Terminal-Bench Hard | 6.8%#240 | Artificial Analysis |
| τ²-Bench Telecom (AA)not in index | 19.0%#329 | Artificial Analysis |
Maths
| Benchmark | default | Source |
|---|---|---|
| OTIS Mock AIME | 9.7%#210 | Epoch AI Benchmarking Hub |
| MATH Level 5 | 49.8%#53 | Epoch AI Benchmarking Hub |
Knowledge
| Benchmark | default | Source |
|---|---|---|
| AA-Omniscience | -17.1#170 | Artificial Analysis |
Instruction following
| Benchmark | default | Source |
|---|---|---|
| IFBench | 39.0%#262 | Artificial Analysis |
Human preference
| Benchmark | default | Source |
|---|---|---|
| LMArena Text | 1335#198 | LMArena |
Long context
| Benchmark | default | Source |
|---|---|---|
| AA-LCR | 25.3%#339 | Artificial Analysis |
Composite
| Benchmark | default | Source |
|---|---|---|
| Epoch Capabilities Indexnot in index | 128.8#138 | Epoch AI Benchmarking Hub |
| AA Intelligence Index | 7.3#405 | Artificial Analysis |
See also
Data as of 9 Sept 2026. Compare with another model.