BenchLeader
AmazonAuto-detected

Nova 2.0 Pro Preview

Best configuration ranks #251 of 610 on the BenchLeader Index at 52.3 ±7.9 (medium reasoning effort). Last measured 2 Dec 2025stale: no new result in six months. Released 2 Dec 2025.

Blended price
$3.44/M
$1.25 in · $10.00 out
Output speed
120 tok/s
First answer
29 s
first token 1.03 s
Context
256k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index52
  2. Coding48
  3. Agents & tools55
  4. Knowledge41
  5. Instruction following76
  6. Multimodal49
  7. Long context58
  8. Composite47

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning104 tok/s1.03 s$0.0035Coding 28
low51.3#268120 tok/s26 s$0.0035Agents & tools 49 · Coding 42 · Composite 45 · Instruction following 77 · Knowledge 45 · Long context 59 · Multimodal 47
mediumbest52.3#251120 tok/s29 s$0.0035Agents & tools 55 · Coding 48 · Composite 47 · Instruction following 76 · Knowledge 41 · Long context 58 · Multimodal 49
default46.7#365104 tok/s1.03 s$0.0035Agents & tools 48 · Composite 42 · Instruction following 53 · Knowledge 41 · Long context 40

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowmediumdefaultSource
GPQA Diamond (AA)not in index75.0%#24378.5%#20363.6%#335Artificial Analysis
Humanity's Last Exam (AA)not in index5.2%#3659.4%#2683.9%#471Artificial Analysis

Coding

Benchmarkno reasoninglowmediumdefaultSource
SciCode28.1%#14338.7%#11842.7%#96SciCode

Agents & tools

Benchmarkno reasoninglowmediumdefaultSource
Terminal-Bench Hard17.4%#17324.2%#13716.7%#179Artificial Analysis
τ²-Bench Telecom (AA)not in index90.6%#6292.7%#5171.6%#146Artificial Analysis

Knowledge

Benchmarkno reasoninglowmediumdefaultSource
AA-Omniscience-40.5#256-49.0#310-49.5#313Artificial Analysis

Instruction following

Benchmarkno reasoninglowmediumdefaultSource
IFBench79.6%#1079.0%#1252.0%#160Artificial Analysis

Multimodal

Benchmarkno reasoninglowmediumdefaultSource
MMMU-Pro62.7%#16064.5%#151Artificial Analysis

Long context

Benchmarkno reasoninglowmediumdefaultSource
AA-LCR64.7%#19564.0%#19930.0%#323Artificial Analysis

Composite

Benchmarkno reasoninglowmediumdefaultSource
AA Intelligence Index12.8#25614.2#23410.0#312Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.003531.6 s
Summarise a 30-page report12,000 / 600$0.02134.1 s
Code edit6,000 / 1,500$0.02241.6 s
Agentic coding session60,000 / 4,000$0.1151.0 min
Structured extraction2,000 / 200$0.004530.8 s

See also

Data as of 9 Sept 2026. Compare these configurations.