BenchLeader

Granite 4.0 H Small

Best configuration ranks #512 of 610 on the BenchLeader Index at 40.3 ±1.9. Released 22 Sept 2025.

Blended price
$0.108/M
$0.060 in · $0.250 out
Output speed
22 tok/s
First answer
21 s
first token 21 s
Context
128k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index40
  2. Agents & tools36
  3. Knowledge35
  4. Instruction following35
  5. Long context31
  6. Composite37

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkdefaultSource
GPQA Diamond (AA)not in index41.6%#447Artificial Analysis
Humanity's Last Exam (AA)not in index3.8%#479Artificial Analysis

Agents & tools

Knowledge

BenchmarkdefaultSource
AA-Omniscience-61.0#394Artificial Analysis

Instruction following

BenchmarkdefaultSource
IFBench31.5%#333Artificial Analysis

Long context

BenchmarkdefaultSource
AA-LCR11.3%#390Artificial Analysis

Composite

BenchmarkdefaultSource
AA Intelligence Index6.0#457Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.000134.5 s
Summarise a 30-page report12,000 / 600$0.000948.0 s
Code edit6,000 / 1,500$0.00071.5 min
Agentic coding session60,000 / 4,000$0.00463.4 min
Structured extraction2,000 / 200$0.000230.0 s

See also

Data as of 9 Sept 2026. Compare with another model.