BenchLeader
GoogleOpen weightsReasoning modelAuto-detected

DiffusionGemma 26B A4B

Best configuration ranks #369 of 610 on the BenchLeader Index at 46.5 ±8.2.

Blended price
Output speed
First answer
Context
256k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index47
  2. Knowledge36
  3. Instruction following59
  4. Multimodal51
  5. Long context35
  6. Composite41

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkdefaultSource
GPQA Diamond (AA)not in index66.9%#312Artificial Analysis
Humanity's Last Exam (AA)not in index10.8%#244Artificial Analysis

Knowledge

BenchmarkdefaultSource
AA-Omniscience-59.2#384Artificial Analysis

Instruction following

BenchmarkdefaultSource
IFBench59.5%#126Artificial Analysis

Multimodal

BenchmarkdefaultSource
MMMU-Pro66.5%#139Artificial Analysis

Long context

BenchmarkdefaultSource
AA-LCR19.7%#365Artificial Analysis

Composite

BenchmarkdefaultSource
AA Intelligence Index9.5#326Artificial Analysis

See also

Data as of 9 Sept 2026. Compare with another model.