BenchLeader
GoogleReasoning model

Gemini 2.5 Flash-Lite

Best configuration ranks #363 of 610 on the BenchLeader Index at 46.7 ±4.8 (thinking reasoning effort). Last measured 2 Sept 2026. Released 17 Jun 2025.

Blended price
$0.175/M
$0.100 in · $0.400 out
Output speed
378 tok/s
First answer
25 s
first token 0.30 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index47
  2. Reasoning54
  3. Coding46
  4. Agents & tools38
  5. Knowledge42
  6. Instruction following51
  7. Human preference55
  8. Multimodal47
  9. Long context40
  10. Composite40

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
thinkingbest46.7#363378 tok/s25 s$0.0002Agents & tools 38 · Coding 46 · Composite 40 · Human preference 55 · Instruction following 51 · Knowledge 42 · Long context 40 · Multimodal 47 · Reasoning 54
default41.6#481290 tok/s0.30 s$0.0002Agents & tools 44 · Coding 40 · Composite 37 · Instruction following 35 · Knowledge 36 · Long context 41 · Multimodal 38 · Reasoning 39

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkthinkingdefaultSource
LMArena Hard Prompts1382#164LMArena
GPQA Diamond (AA)not in index62.5%#34347.4%#417Artificial Analysis
Humanity's Last Exam (AA)not in index6.8%#3083.7%#491Artificial Analysis
Kagi LLM Benchmark40.5%#102Kagi LLM Benchmark

Coding

BenchmarkthinkingdefaultSource
WeirdML35.2%#12035.2%#120WeirdML
LMArena Coding1383#190LMArena

Agents & tools

BenchmarkthinkingdefaultSource
Terminal-Bench Hard4.5%#2702.3%#304Artificial Analysis
τ²-Bench Telecom (AA)not in index18.4%#33219.0%#329Artificial Analysis
BFCL Overall36.9%#40Berkeley Function Calling Leaderboard

Knowledge

BenchmarkthinkingdefaultSource
AA-Omniscience-45.6#287-58.8#380Artificial Analysis

Instruction following

BenchmarkthinkingdefaultSource
IFBench49.9%#17231.5%#333Artificial Analysis

Human preference

BenchmarkthinkingdefaultSource
LMArena Text1375#156LMArena

Multimodal

BenchmarkthinkingdefaultSource
LMArena Vision1187#80LMArena
MMMU-Pro58.2%#18554.0%#197Artificial Analysis

Long context

BenchmarkthinkingdefaultSource
Fiction.LiveBench 120k21.9%#33Fiction.live
AA-LCR55.7%#23232.0%#315Artificial Analysis

Composite

BenchmarkthinkingdefaultSource
Epoch Capabilities Indexnot in index133.9#122Epoch AI Benchmarking Hub
AA Intelligence Index8.6#3586.7#431Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Google Vertex (EU)154 tok/s0.41 s$0.100$0.4001.0M
Google Vertex90 tok/s1.02 s$0.100$0.4001.0M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.010 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0002$0.000125.8 s
Summarise a 30-page report12,000 / 600$0.0014$0.000626.5 s
Code edit6,000 / 1,500$0.0012$0.000828.9 s
Agentic coding session60,000 / 4,000$0.0076$0.003635.5 s
Structured extraction2,000 / 200$0.0003$0.000125.5 s

See also

Data as of 9 Sept 2026. Compare these configurations.