AllenaiAuto-detected
Olmo 3 32B
Best configuration ranks #438 of 610 on the BenchLeader Index at 43.4 ±5.9 (thinking reasoning effort). Last measured 2 Sept 2026.
- Blended price
- –
- Output speed
- –
- First answer
- –
- Context
- 66k
- Overall index43
- Reasoning48
- Coding49
- Agents & tools35
- Knowledge37
- Instruction following50
- Human preference47
- Long context25
- Composite37
Benchmark results
One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.
Reasoning
| Benchmark | thinking | Source |
|---|---|---|
| LMArena Hard Prompts | 1327#213 | LMArena |
| GPQA Diamond (AA)not in index | 61.0%#351 | Artificial Analysis |
| Humanity's Last Exam (AA)not in index | 6.4%#317 | Artificial Analysis |
Coding
| Benchmark | thinking | Source |
|---|---|---|
| LMArena Coding | 1364#206 | LMArena |
Agents & tools
| Benchmark | thinking | Source |
|---|---|---|
| Terminal-Bench Hard | 1.5%#321 | Artificial Analysis |
| τ²-Bench Telecom (AA)not in index | 0.0%#375 | Artificial Analysis |
Knowledge
| Benchmark | thinking | Source |
|---|---|---|
| AA-Omniscience | -58.1#374 | Artificial Analysis |
Instruction following
| Benchmark | thinking | Source |
|---|---|---|
| IFBench | 49.1%#178 | Artificial Analysis |
Human preference
| Benchmark | thinking | Source |
|---|---|---|
| LMArena Text | 1307#226 | LMArena |
Long context
| Benchmark | thinking | Source |
|---|---|---|
| AA-LCR | 0.0%#419 | Artificial Analysis |
Composite
| Benchmark | thinking | Source |
|---|---|---|
| AA Intelligence Index | 6.5#442 | Artificial Analysis |
See also
Data as of 9 Sept 2026. Compare with another model.