BenchLeader
AllenaiAuto-detected

Olmo 3 32B

Best configuration ranks #438 of 610 on the BenchLeader Index at 43.4 ±5.9 (thinking reasoning effort). Last measured 2 Sept 2026.

Blended price
Output speed
First answer
Context
66k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index43
  2. Reasoning48
  3. Coding49
  4. Agents & tools35
  5. Knowledge37
  6. Instruction following50
  7. Human preference47
  8. Long context25
  9. Composite37

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkthinkingSource
LMArena Hard Prompts1327#213LMArena
GPQA Diamond (AA)not in index61.0%#351Artificial Analysis
Humanity's Last Exam (AA)not in index6.4%#317Artificial Analysis

Coding

BenchmarkthinkingSource
LMArena Coding1364#206LMArena

Agents & tools

Knowledge

BenchmarkthinkingSource
AA-Omniscience-58.1#374Artificial Analysis

Instruction following

BenchmarkthinkingSource
IFBench49.1%#178Artificial Analysis

Human preference

BenchmarkthinkingSource
LMArena Text1307#226LMArena

Long context

BenchmarkthinkingSource
AA-LCR0.0%#419Artificial Analysis

Composite

BenchmarkthinkingSource
AA Intelligence Index6.5#442Artificial Analysis

See also

Data as of 9 Sept 2026. Compare with another model.