BenchLeader
AmazonReasoning modelAuto-detected

Nova 2.0 Lite (high)

Best configuration ranks #282 of 610 on the BenchLeader Index at 50.6 ±7.7 (thinking reasoning effort). Released 29 Oct 2025.

Blended price
$0.850/M
$0.300 in · $2.50 out
Output speed
151 tok/s
First answer
30 s
first token 1.00 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index51
  2. Agents & tools48
  3. Knowledge38
  4. Instruction following69
  5. Multimodal48
  6. Long context56
  7. Composite46

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
low47.2#353162 tok/s24 s$0.0009Agents & tools 37 · Composite 44 · Instruction following 61 · Knowledge 38 · Long context 53 · Multimodal 42
medium50.0#299151 tok/s28 s$0.0009Agents & tools 49 · Composite 45 · Instruction following 67 · Knowledge 37 · Long context 56 · Multimodal 47
thinkingbest50.6#282151 tok/s30 s$0.0009Agents & tools 48 · Composite 46 · Instruction following 69 · Knowledge 38 · Long context 56 · Multimodal 48
default41.7#479158 tok/s1.00 s$0.0009Agents & tools 40 · Composite 40 · Instruction following 43 · Knowledge 36 · Long context 34 · Multimodal 33

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarklowmediumthinkingdefaultSource
GPQA Diamond (AA)not in index69.8%#28976.8%#22081.1%#17860.3%#355Artificial Analysis
Humanity's Last Exam (AA)not in index4.0%#4639.0%#27111.6%#2272.9%#517Artificial Analysis

Agents & tools

BenchmarklowmediumthinkingdefaultSource
Terminal-Bench Hard3.8%#28317.4%#17316.7%#1796.8%#240Artificial Analysis
τ²-Bench Telecom (AA)not in index71.9%#14575.7%#13272.8%#14262.0%#172Artificial Analysis

Knowledge

BenchmarklowmediumthinkingdefaultSource
AA-Omniscience-54.1#351-57.0#366-55.0#361-59.8#388Artificial Analysis

Instruction following

BenchmarklowmediumthinkingdefaultSource
IFBench61.2%#11968.5%#7970.8%#6440.5%#247Artificial Analysis

Multimodal

BenchmarklowmediumthinkingdefaultSource
MMMU-Pro58.0%#18762.5%#16163.8%#15549.0%#208Artificial Analysis

Long context

BenchmarklowmediumthinkingdefaultSource
AA-LCR54.3%#23860.0%#21960.3%#21618.7%#369Artificial Analysis

Composite

BenchmarklowmediumthinkingdefaultSource
AA Intelligence Index11.8#27812.5#26113.4#2458.7#353Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.000932.2 s
Summarise a 30-page report12,000 / 600$0.005134.2 s
Code edit6,000 / 1,500$0.005640.1 s
Agentic coding session60,000 / 4,000$0.02856.6 s
Structured extraction2,000 / 200$0.001131.5 s

See also

Data as of 9 Sept 2026. Compare these configurations.