BenchLeader
GoogleReasoning model

Gemini 2.5 Flash

Best configuration ranks #312 of 610 on the BenchLeader Index at 48.8 ±3.6. Last measured 2 Sept 2026. Released 17 Apr 2025.

Blended price
$0.850/M
$0.300 in · $2.50 out
Output speed
193 tok/s
First answer
0.43 s
first token 0.86 s
Context
1.0M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index49
  2. Reasoning42
  3. Coding43
  4. Agents & tools44
  5. Maths56
  6. Knowledge47
  7. Instruction following42
  8. Human preference60
  9. Multimodal58
  10. Long context55
  11. Composite41

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
thinking48.1#334211 tok/s17 s$0.0009Agents & tools 45 · Coding 35 · Composite 46 · Instruction following 51 · Knowledge 52 · Long context 59 · Multimodal 53 · Reasoning 40
defaultbest48.8#312193 tok/s0.43 s$0.0009Agents & tools 44 · Coding 43 · Composite 41 · Human preference 60 · Instruction following 42 · Knowledge 47 · Long context 55 · Maths 56 · Multimodal 58 · Reasoning 42

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkthinkingdefaultSource
Humanity's Last Exam12.1%#24Scale AI / CAIS
SimpleBench41.2%#58SimpleBench
LMArena Hard Prompts1420#130LMArena
GPQA Diamond (AA)not in index79.0%#19768.3%#300Artificial Analysis
Humanity's Last Exam (AA)not in index12.1%#2184.7%#400Artificial Analysis
GPQA Diamond (Vals)not in index57.6%#110Vals AI
Kagi LLM Benchmark56.8%#6044.1%#96Kagi LLM Benchmark
ARC-AGI-133.3%#14033.3%#140ARC Prize
ARC-AGI-22.5%#1391.7%#148ARC Prize

Coding

BenchmarkthinkingdefaultSource
WeirdML41.0%#9741.0%#97WeirdML
LMArena Coding1424#157LMArena
LiveCodeBench46.9%#11456.9%#106Vals AI
IOI2.6%#53Vals AI
Aider Polyglot55.1%#18Aider polyglot leaderboard
SWE-bench Verified (bash only)28.7%#37SWE-bench
SWE-bench Verified (any scaffold)not in index28.7%#56SWE-bench

Agents & tools

Maths

BenchmarkthinkingdefaultSource
OTIS Mock AIME73.1%#119Epoch AI Benchmarking Hub

Knowledge

BenchmarkthinkingdefaultSource
AA-Omniscience-29.8#213-42.6#269Artificial Analysis
LegalBench83.8%#43Vals AI
CorpFin54.2%#92Vals AI
TaxEval70.5%#8871.2%#80Vals AI
MedQA91.0%#4286.7%#57Vals AI
PRBench Finance38.4%#27Scale AI SEAL
PRBench Legal41.0%#20Scale AI SEAL

Instruction following

BenchmarkthinkingdefaultSource
IFBench50.3%#17039.0%#263Artificial Analysis

Human preference

BenchmarkthinkingdefaultSource
LMArena Text1410#122LMArena

Multimodal

BenchmarkthinkingdefaultSource
LMArena Vision1236#58LMArena
MMMU-Pro69.1%#12665.5%#141Artificial Analysis
VISTA49.1%#13Scale AI SEAL
MMMU (validation)79.7%#6MMMU

Long context

BenchmarkthinkingdefaultSource
Fiction.LiveBench 120k68.8%#8Fiction.live
AA-LCR65.3%#19149.9%#255Artificial Analysis

Composite

BenchmarkthinkingdefaultSource
Epoch Capabilities Indexnot in index143#82Epoch AI Benchmarking Hub
AA Intelligence Index13.1#2539.8#318Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Google Vertex (EU)96 tok/s0.50 s$0.300$2.501.0M
Google Vertex (Global) (ZDR)73 tok/s0.86 s$0.300$2.501.0M
Google Vertex (Global) (ZDR)42 tok/s1.26 s$0.540$4.501.0M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.030 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0009$0.00082.0 s
Summarise a 30-page report12,000 / 600$0.0051$0.00273.5 s
Code edit6,000 / 1,500$0.0056$0.00438.2 s
Agentic coding session60,000 / 4,000$0.028$0.01621.2 s
Structured extraction2,000 / 200$0.0011$0.00071.5 s

See also

Data as of 9 Sept 2026. Compare these configurations.