BenchLeader
GoogleAuto-detected

Gemini 2.5 Flash 09

Best configuration ranks #246 of 610 on the BenchLeader Index at 52.5 ±2.4 (thinking reasoning effort). Last measured 1 Sept 2026. Released 25 Sept 2025.

Blended price
$0.850/M
$0.300 in · $2.50 out
Output speed
First answer
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index53
  2. Coding54
  3. Agents & tools48
  4. Maths48
  5. Knowledge53
  6. Instruction following53
  7. Multimodal58
  8. Long context62
  9. Composite49

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
thinkingbest52.5#246$0.0009Agents & tools 48 · Coding 54 · Composite 49 · Instruction following 53 · Knowledge 53 · Long context 62 · Maths 48 · Multimodal 58
default50.7#279$0.0009Agents & tools 36 · Coding 52 · Composite 45 · Human preference 59 · Instruction following 45 · Knowledge 52 · Long context 56 · Maths 47 · Multimodal 57 · Reasoning 59

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkthinkingdefaultSource
LMArena Hard Prompts1420#129LMArena
GPQA Diamond (AA)not in index79.3%#19376.6%#224Artificial Analysis
Humanity's Last Exam (AA)not in index13.8%#2058.7%#277Artificial Analysis
GPQA Diamond (Vals)not in index76.5%#7581.6%#62Vals AI

Coding

BenchmarkthinkingdefaultSource
WeirdML41.9%#92WeirdML
LMArena Coding1428#154LMArena
LiveCodeBench76.2%#7675.1%#77Vals AI

Agents & tools

BenchmarkthinkingdefaultSource
Terminal-Bench17.1%#59Terminal-Bench
Terminal-Bench Hard16.7%#17914.4%#189Artificial Analysis
τ²-Bench Telecom (AA)not in index45.6%#20328.4%#260Artificial Analysis

Maths

BenchmarkthinkingdefaultSource
AIME (Vals)51.5%#5949.8%#60Vals AI
MGSM89.8%#4189.8%#40Vals AI

Knowledge

BenchmarkthinkingdefaultSource
AA-Omniscience-36.1#238-39.9#254Artificial Analysis
MMLU-Pro83.7%#6983.7%#68Vals AI
LegalBench82.6%#5982.5%#62Vals AI
CorpFin59.8%#7159.0%#77Vals AI
TaxEval72.4%#6272.7%#58Vals AI
MedQA91.2%#4091.4%#37Vals AI

Instruction following

BenchmarkthinkingdefaultSource
IFBench52.3%#15943.5%#218Artificial Analysis

Human preference

BenchmarkthinkingdefaultSource
LMArena Text1404#127LMArena

Multimodal

BenchmarkthinkingdefaultSource
LMArena Vision1253#47LMArena
MMMU-Pro73.1%#9870.2%#115Artificial Analysis

Long context

BenchmarkthinkingdefaultSource
AA-LCR71.0%#15260.0%#219Artificial Analysis

Composite

BenchmarkthinkingdefaultSource
AA Intelligence Index15.5#21412.4#267Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0009
Summarise a 30-page report12,000 / 600$0.0051
Code edit6,000 / 1,500$0.0056
Agentic coding session60,000 / 4,000$0.028
Structured extraction2,000 / 200$0.0011

See also

Data as of 9 Sept 2026. Compare these configurations.