BenchLeader
GoogleAuto-detected

Gemini 2.5 Flash Lite 09

Best configuration ranks #314 of 610 on the BenchLeader Index at 48.8 ±3.1 (thinking reasoning effort). Last measured 1 Sept 2026.

Blended price
$0.175/M
$0.100 in · $0.400 out
Output speed
First answer
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index49
  2. Coding50
  3. Agents & tools45
  4. Maths44
  5. Knowledge47
  6. Instruction following53
  7. Multimodal49
  8. Long context59
  9. Composite42

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
thinkingbest48.8#314$0.0002Agents & tools 45 · Coding 50 · Composite 42 · Instruction following 53 · Knowledge 47 · Long context 59 · Maths 44 · Multimodal 49
default45.7#384$0.0002Agents & tools 40 · Coding 47 · Composite 41 · Instruction following 44 · Knowledge 46 · Long context 51 · Maths 42 · Multimodal 48

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkthinkingdefaultSource
GPQA Diamond (AA)not in index70.9%#28165.0%#328Artificial Analysis
Humanity's Last Exam (AA)not in index7.0%#3045.1%#369Artificial Analysis
GPQA Diamond (Vals)not in index70.2%#8964.1%#101Vals AI

Coding

BenchmarkthinkingdefaultSource
LiveCodeBench71.4%#8367.7%#92Vals AI

Agents & tools

BenchmarkthinkingdefaultSource
Terminal-Bench Hard12.9%#1997.6%#232Artificial Analysis
τ²-Bench Telecom (AA)not in index30.7%#24930.4%#251Artificial Analysis

Maths

BenchmarkthinkingdefaultSource
AIME (Vals)42.1%#6626.3%#75Vals AI
MGSM88.4%#5089.5%#43Vals AI

Knowledge

BenchmarkthinkingdefaultSource
AA-Omniscience-54.0#349-42.6#269Artificial Analysis
MMLU-Pro79.1%#9578.6%#98Vals AI
LegalBench82.0%#7079.0%#92Vals AI
CorpFin57.6%#8556.3%#88Vals AI
TaxEval64.7%#11366.2%#111Vals AI
MedQA88.9%#5180.3%#70Vals AI

Instruction following

BenchmarkthinkingdefaultSource
IFBench52.6%#15741.8%#237Artificial Analysis

Multimodal

BenchmarkthinkingdefaultSource
MMMU-Pro65.0%#14563.4%#156Artificial Analysis

Long context

BenchmarkthinkingdefaultSource
AA-LCR64.7%#19549.7%#256Artificial Analysis

Composite

BenchmarkthinkingdefaultSource
AA Intelligence Index10.3#3079.3#332Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0002
Summarise a 30-page report12,000 / 600$0.0014
Code edit6,000 / 1,500$0.0012
Agentic coding session60,000 / 4,000$0.0076
Structured extraction2,000 / 200$0.0003

See also

Data as of 9 Sept 2026. Compare these configurations.