BenchLeader
GoogleReasoning modelAuto-detected

Gemini 3.5 Flash Lite

Best configuration ranks #221 of 610 on the BenchLeader Index at 53.5 ±8.1. Last measured 8 Sept 2026. Released 21 Jul 2026.

Blended price
$0.850/M
$0.300 in · $2.50 out
Output speed
369 tok/s
First answer
8.01 s
first token 1.07 s
Context
1.0M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index54
  2. Reasoning65
  3. Coding53
  4. Agents & tools18
  5. Maths38
  6. Knowledge67
  7. Human preference65
  8. Multimodal63
  9. Long context65
  10. Composite58

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
minimal369 tok/s8.01 s$0.0009Maths 46 · Reasoning 41
low369 tok/s8.01 s$0.0009Maths 50 · Reasoning 42
medium369 tok/s8.01 s$0.0009Reasoning 37
high44.0#426369 tok/s8.01 s$0.0009Agents & tools 39 · Coding 52 · Composite 18 · Knowledge 57 · Maths 42 · Reasoning 49
defaultbest53.5#221369 tok/s8.01 s$0.0009Agents & tools 18 · Coding 53 · Composite 58 · Human preference 65 · Knowledge 67 · Long context 65 · Maths 38 · Multimodal 63 · Reasoning 65

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkminimallowmediumhighdefaultSourceTrend
GPQA Diamond74.2%#14175.8%#13183.3%#91Epoch AI Benchmarking Hub
LMArena Hard Prompts1476#60LMArena
LiveBench Reasoningnot in index60.2%#52LiveBench
GPQA Diamond (AA)not in index83.8%#151Artificial Analysis
Humanity's Last Exam (AA)not in index18.8%#169Artificial Analysis
GPQA Diamond (Vals)not in index83.8%#54Vals AI
ARC-AGI-17.5%#17017.0%#15932.3%#14553.5%#120ARC Prize
ARC-AGI-20.8%#1641.5%#1515.3%#11810.3%#105ARC Prize

Coding

BenchmarkminimallowmediumhighdefaultSourceTrend
SciCode40.9%#102SciCode
WeirdML39.0%#108WeirdML
LMArena Coding1503#59LMArena
LMArena WebDev1449#53LMArena
LiveBench Codingnot in index76.1%#33LiveBench
SciCode (AA)not in index41.3%#110Artificial Analysis
LiveCodeBench79.0%#72Vals AI
IOI26.2%#17Vals AI
SWE-bench (Vals)not in index75.0%#42Vals AI

Agents & tools

BenchmarkminimallowmediumhighdefaultSourceTrend
LMArena Agent-13.4#40LMArena
LiveBench Agentic Codingnot in index45.3%#41LiveBench
Terminal-Bench 2.1 (Vals)50.2%#48Vals AI

Maths

BenchmarkminimallowmediumhighdefaultSourceTrend
FrontierMath Tiers 1–326.0%#64Epoch AI Benchmarking Hub
FrontierMath Tier 40.0%#57Epoch AI Benchmarking Hub
OTIS Mock AIME51.1%#16560.0%#14671.1%#120Epoch AI Benchmarking Hub
ProofBench13.0%#47Vals AI
LiveBench Mathematicsnot in index73.7%#52LiveBench

Knowledge

BenchmarkminimallowmediumhighdefaultSourceTrend
LiveBench Data Analysisnot in index53.3%#52LiveBench
AA-Omniscience5.2#72Artificial Analysis
MMLU-Pro85.8%#51Vals AI
LegalBench84.1%#36Vals AI
CorpFin60.7%#65Vals AI
TaxEval72.6%#60Vals AI

Instruction following

BenchmarkminimallowmediumhighdefaultSourceTrend
LiveBench Languagenot in index71.8%#47LiveBench

Human preference

BenchmarkminimallowmediumhighdefaultSourceTrend
LMArena Text1457#56LMArena

Multimodal

BenchmarkminimallowmediumhighdefaultSourceTrend
LMArena Vision1269#34LMArena
MMMU-Pro79.0%#45Artificial Analysis

Long context

BenchmarkminimallowmediumhighdefaultSourceTrend
AA-LCR76.0%#102Artificial Analysis

Composite

BenchmarkminimallowmediumhighdefaultSourceTrend
Epoch Capabilities Indexnot in index145.1#71Epoch AI Benchmarking Hub
LiveBench63.9%#51LiveBench
AA Intelligence Index22.7#142Artificial Analysis
Vals Indexnot in index36.7#38Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Google Vertex Flex169 tok/s6.74 s$0.150$1.251.0M
Google AI Studio72 tok/s0.77 s$0.300$2.501.0M
Google Vertex68 tok/s0.61 s$0.300$2.501.0M
Google AI Studio Priority50 tok/s1.12 s$0.540$4.501.0M
Google Vertex Priority33 tok/s1.07 s$0.540$4.501.0M
Google Vertex (EU)23 tok/s1.49 s$0.330$2.751.0M
Google AI Studio Flex17 tok/s0.70 s$0.150$1.251.0M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.030 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0009$0.00088.8 s
Summarise a 30-page report12,000 / 600$0.0051$0.00279.6 s
Code edit6,000 / 1,500$0.0056$0.004312.1 s
Agentic coding session60,000 / 4,000$0.028$0.01618.8 s
Structured extraction2,000 / 200$0.0011$0.00078.6 s

See also

Data as of 9 Sept 2026. Compare these configurations.