BenchLeader
NVIDIAOpen weightsAuto-detected

Llama Nemotron Super 49B v1.5

Best configuration ranks #458 of 610 on the BenchLeader Index at 42.6 ±1.8. Released 25 Jul 2025.

Blended price
$0.400/M
$0.400 in · $0.400 out
Output speed
35 tok/s
First answer
12 s
first token 12 s
Context
128k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index43
  2. Agents & tools37
  3. Knowledge42
  4. Instruction following36
  5. Long context38
  6. Composite38

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkdefaultSource
GPQA Diamond (AA)not in index48.1%#415Artificial Analysis
Humanity's Last Exam (AA)not in index4.3%#429Artificial Analysis

Agents & tools

Knowledge

BenchmarkdefaultSource
AA-Omniscience-46.2#289Artificial Analysis

Instruction following

BenchmarkdefaultSource
IFBench32.9%#320Artificial Analysis

Long context

BenchmarkdefaultSource
AA-LCR24.7%#341Artificial Analysis

Composite

BenchmarkdefaultSource
AA Intelligence Index7.4#401Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.000320.4 s
Summarise a 30-page report12,000 / 600$0.005028.9 s
Code edit6,000 / 1,500$0.003054.7 s
Agentic coding session60,000 / 4,000$0.0262.1 min
Structured extraction2,000 / 200$0.000917.5 s

See also

Data as of 9 Sept 2026. Compare with another model.