BenchLeader
GoogleAuto-detected

Gemini 2.5 Flash 09 2025

Gemini 2.5 Flash 09 2025 is a Google proprietary model, released 25 Sept 2025. Its best configuration (thinking reasoning effort) ranks #248 of 372 on the BenchLeader Index at 52.6 ±2.4, in the lower half. It scores highest in long context (62) and lowest in agents & tools (48). At $0.850 per million tokens blended it is mid-priced. It has been measured at 2 reasoning-effort settings; tables show the best-scoring one. Last measured 10 Sept 2026.

Blended price
$0.850/M
$0.300 in · $2.50 out
Output speed
First answer
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index53
  2. Coding55
  3. Agents & tools48
  4. Maths48
  5. Knowledge53
  6. Instruction following53
  7. Multimodal57
  8. Long context62
  9. Composite49

Versions

Google has shipped 12 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
Gemini 3.8 Flashmedium2 Sept 202663.6#45
Gemini 3.7 Flashmedium13 Aug 202664.5#34
Gemini 3.6 Flash21 Jul 202660.8#80
Gemini 3.5 Flashmedium19 May 202664.1#37
Gemini 2.5 Flash 09 2025thinkingthis page25 Sept 202552.6#248
Gemini 2.5 Flash17 Apr 202549.2#320
Gemini 2.0 Flashthinking5 Feb 202545.2#404
Gemini 1.5 Flash 00224 Sept 202439.0#557
Gemini 1.5 Flash 00123 May 202436.5#594
Gemini 3 Flashthinking61.6#67
Gemini 1.5 Flash
Gemini Flash 2.0

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
thinkingbest52.6#248$0.0009Agents & tools 48 · Coding 55 · Composite 49 · Instruction following 53 · Knowledge 53 · Long context 62 · Maths 48 · Multimodal 57
default50.9#283$0.0009Agents & tools 36 · Coding 52 · Composite 45 · Human preference 59 · Instruction following 46 · Knowledge 53 · Long context 56 · Maths 48 · Multimodal 57 · Reasoning 59

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkthinkingdefaultSource
LMArena Hard Prompts1420#130LMArena
GPQA Diamond (AA)not in index79.3%#19676.6%#227Artificial Analysis
Humanity's Last Exam (AA)not in index13.8%#2098.7%#283Artificial Analysis
GPQA Diamond (Vals)not in index76.5%#7581.6%#63Vals AI

Coding

BenchmarkthinkingdefaultSource
WeirdML41.9%#94WeirdML
LMArena Coding1428#157LMArena
LiveCodeBench76.2%#7775.1%#78Vals AI

Agents & tools

BenchmarkthinkingdefaultSource
Terminal-Bench17.1%#59Terminal-Bench
Terminal-Bench Hard16.7%#17914.4%#190Artificial Analysis
τ²-Bench Telecom (AA)not in index45.6%#20528.4%#260Artificial Analysis

Maths

BenchmarkthinkingdefaultSource
AIME (Vals)51.5%#5849.8%#59Vals AI
MGSM89.8%#4089.8%#39Vals AI

Knowledge

BenchmarkthinkingdefaultSource
AA-Omniscience-36.1#244-39.9#260Artificial Analysis
MMLU-Pro83.7%#7083.7%#69Vals AI
LegalBench82.6%#6082.5%#63Vals AI
CorpFin59.8%#7259.0%#78Vals AI
TaxEval72.4%#6172.7%#57Vals AI
MedQA91.2%#4091.4%#37Vals AI

Instruction following

BenchmarkthinkingdefaultSource
IFBench52.3%#16043.5%#218Artificial Analysis

Human preference

BenchmarkthinkingdefaultSource
LMArena Text1404#128LMArena

Multimodal

BenchmarkthinkingdefaultSource
LMArena Vision1253#47LMArena
MMMU-Pro73.1%#10070.2%#117Artificial Analysis

Long context

BenchmarkthinkingdefaultSource
AA-LCR71.0%#15660.0%#223Artificial Analysis

Composite

BenchmarkthinkingdefaultSource
AA Intelligence Index15.5#21912.4#271Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0009
Summarise a 30-page report12,000 / 600$0.0051
Code edit6,000 / 1,500$0.0056
Agentic coding session60,000 / 4,000$0.028
Structured extraction2,000 / 200$0.0011

See also

Data as of 10 Sept 2026. Compare these configurations.