BenchLeader
AmazonReasoning modelAuto-detected

Nova 2.0 Omni (medium)

Best configuration ranks #323 of 610 on the BenchLeader Index at 48.5 ±8.1 (medium reasoning effort).

Blended price
$0.850/M
$0.300 in · $2.50 out
Output speed
First answer
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index49
  2. Agents & tools38
  3. Knowledge36
  4. Instruction following65
  5. Multimodal46
  6. Long context56
  7. Composite46

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
low47.8#346$0.0009Agents & tools 37 · Composite 43 · Instruction following 61 · Knowledge 40 · Long context 55 · Multimodal 44
mediumbest48.5#323$0.0009Agents & tools 38 · Composite 46 · Instruction following 65 · Knowledge 36 · Long context 56 · Multimodal 46
default41.9#475$0.0009Agents & tools 40 · Composite 39 · Instruction following 43 · Knowledge 34 · Long context 38 · Multimodal 34

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarklowmediumdefaultSource
GPQA Diamond (AA)not in index69.9%#28776.0%#23555.5%#385Artificial Analysis
Humanity's Last Exam (AA)not in index7.0%#3043.9%#474Artificial Analysis

Agents & tools

BenchmarklowmediumdefaultSource
Terminal-Bench Hard3.8%#2834.5%#2706.8%#240Artificial Analysis
τ²-Bench Telecom (AA)not in index67.8%#15880.4%#11944.7%#207Artificial Analysis

Knowledge

BenchmarklowmediumdefaultSource
AA-Omniscience-50.3#315-59.1#382-64.2#418Artificial Analysis

Instruction following

BenchmarklowmediumdefaultSource
IFBench61.8%#11866.2%#9741.1%#245Artificial Analysis

Multimodal

BenchmarklowmediumdefaultSource
MMMU-Pro59.8%#17961.9%#16949.9%#206Artificial Analysis

Long context

BenchmarklowmediumdefaultSource
AA-LCR57.7%#22759.7%#22224.7%#341Artificial Analysis

Composite

BenchmarklowmediumdefaultSource
AA Intelligence Index11.1#29413.7#2398.2#368Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0009
Summarise a 30-page report12,000 / 600$0.0051
Code edit6,000 / 1,500$0.0056
Agentic coding session60,000 / 4,000$0.028
Structured extraction2,000 / 200$0.0011

See also

Data as of 9 Sept 2026. Compare these configurations.