BenchLeader
GoogleAuto-detected

Gemini Pro

Not yet ranked: too few independent results so far.

Blended price
$3.44/M
$1.25 in · $10.00 out
Output speed
First answer
Context
1.0M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Reasoning37
  2. Coding36
  3. Human preference37

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkdefaultSource
LMArena Hard Prompts1230#277LMArena

Coding

BenchmarkdefaultSource
LMArena Coding1249#281LMArena

Human preference

BenchmarkdefaultSource
LMArena Text1223#278LMArena

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0035
Summarise a 30-page report12,000 / 600$0.021
Code edit6,000 / 1,500$0.022
Agentic coding session60,000 / 4,000$0.115
Structured extraction2,000 / 200$0.0045

See also

Data as of 9 Sept 2026. Compare with another model.