BenchLeader
Allen Institute for AIOpen weightsAuto-detected

Molmo2-8B

Best configuration ranks #583 of 610 on the BenchLeader Index at 36.3 ±4.0.

Blended price
Output speed
First answer
Context
37k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index36
  2. Agents & tools34
  3. Knowledge31
  4. Instruction following31
  5. Multimodal21
  6. Long context25
  7. Composite35

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkdefaultSource
GPQA Diamond (AA)not in index42.5%#443Artificial Analysis
Humanity's Last Exam (AA)not in index4.3%#434Artificial Analysis

Agents & tools

Knowledge

BenchmarkdefaultSource
AA-Omniscience-69.2#432Artificial Analysis

Instruction following

BenchmarkdefaultSource
IFBench26.9%#362Artificial Analysis

Multimodal

BenchmarkdefaultSource
MMMU-Pro37.5%#236Artificial Analysis

Long context

BenchmarkdefaultSource
AA-LCR0.0%#419Artificial Analysis

Composite

BenchmarkdefaultSource
AA Intelligence Index5#522Artificial Analysis

See also

Data as of 9 Sept 2026. Compare with another model.