BenchLeader
GoogleReasoning modelAuto-detected

Gemini 3.6 Flash

Best configuration ranks #79 of 610 on the BenchLeader Index at 60.7 ±5.8. Last measured 21 Jul 2026. Released 21 Jul 2026.

Blended price
$1.50/M
$0.750 in · $3.75 out
Output speed
190 tok/s
First answer
16 s
first token 1.34 s
Context
1.0M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index61
  2. Coding50
  3. Maths55
  4. Knowledge75
  5. Multimodal68
  6. Long context67
  7. Composite72

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
minimal190 tok/s16 s$0.0014Maths 59 · Reasoning 47
low190 tok/s16 s$0.0014Maths 60 · Reasoning 56
medium190 tok/s16 s$0.0014Reasoning 58
high58.6#119190 tok/s16 s$0.0014Agents & tools 49 · Coding 62 · Composite 47 · Human preference 68 · Knowledge 64 · Maths 57 · Multimodal 66 · Reasoning 65
defaultbest60.7#79190 tok/s16 s$0.0014Coding 50 · Composite 72 · Knowledge 75 · Long context 67 · Maths 55 · Multimodal 68

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkminimallowmediumhighdefaultSourceTrend
GPQA Diamond85.9%#7286.4%#6694.1%#6Epoch AI Benchmarking Hub
LMArena Hard Prompts1500#24LMArena
LiveBench Reasoningnot in index85.2%#29LiveBench
GPQA Diamond (AA)not in index92.8%#28Artificial Analysis
Humanity's Last Exam (AA)not in index40.8%#51Artificial Analysis
GPQA Diamond (Vals)not in index93.4%#7Vals AI
ARC-AGI-134.5%#13876.5%#8983.2%#7991.2%#45ARC Prize
ARC-AGI-22.6%#13730.4%#8850.4%#7360.4%#56ARC Prize

Coding

BenchmarkminimallowmediumhighdefaultSourceTrend
SciCode52.7%#47SciCode
WeirdML56.1%#54WeirdML
FrontierCode34.4%#17Cognition
LMArena Coding1518#32LMArena
LMArena WebDev1537#30LMArena
LiveBench Codingnot in index77.9%#25LiveBench
SciCode (AA)not in index53.4%#48Artificial Analysis
LiveCodeBench88.1%#8Vals AI
SWE-bench (Vals)not in index79.6%#25Vals AI

Agents & tools

BenchmarkminimallowmediumhighdefaultSourceTrend
LMArena Agent-5.3#33LMArena
LiveBench Agentic Codingnot in index43.4%#44LiveBench
Terminal-Bench 2.1 (Vals)73.8%#16Vals AI

Maths

BenchmarkminimallowmediumhighdefaultSourceTrend
FrontierMath Tiers 1–359.0%#29Epoch AI Benchmarking Hub
FrontierMath Tier 421.9%#38Epoch AI Benchmarking Hub
OTIS Mock AIME80.0%#9982.2%#9394.2%#41Epoch AI Benchmarking Hub
ProofBench36.0%#25Vals AI
LiveBench Mathematicsnot in index86.4%#37LiveBench
AIME 202696.7%#6MathArena
HMMT February 202689.4%#14MathArena
MathArena Apex26.0%#12MathArena

Knowledge

BenchmarkminimallowmediumhighdefaultSourceTrend
SimpleQA Verified66.2%#9Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index63.0%#47LiveBench
AA-Omniscience22.1#34Artificial Analysis
MMLU-Pro89.3%#13Vals AI
LegalBench86.7%#10Vals AI
CorpFin63.3%#45Vals AI
TaxEval74.9%#25Vals AI

Instruction following

BenchmarkminimallowmediumhighdefaultSourceTrend
LiveBench Languagenot in index83.9%#12LiveBench

Human preference

BenchmarkminimallowmediumhighdefaultSourceTrend
LMArena Text1480#21LMArena

Multimodal

BenchmarkminimallowmediumhighdefaultSourceTrend
LMArena Vision1302#13LMArena
MMMU-Pro83.2%#17Artificial Analysis

Long context

BenchmarkminimallowmediumhighdefaultSourceTrend
AA-LCR80.0%#51Artificial Analysis

Composite

BenchmarkminimallowmediumhighdefaultSourceTrend
Epoch Capabilities Indexnot in index154.3#29Epoch AI Benchmarking Hub
LiveBench73.6%#34LiveBench
AA Intelligence Index34.3#58Artificial Analysis
Vals Indexnot in index55.4#20Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Google AI Studio Priority139 tok/s1.80 s$1.35$6.751.0M
Google Vertex122 tok/s1.55 s$0.750$3.751.0M
Google Vertex Flex110 tok/s7.62 s$0.375$1.881.0M
Google AI Studio Flex105 tok/s0.89 s$0.375$1.881.0M
Google Vertex Priority64 tok/s0.90 s$1.35$6.751.0M
Google AI Studio61 tok/s1.12 s$0.750$3.751.0M

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$2.27$4.54$6.81$9.08Aug 26
input outputnow $0.825 in · $4.13 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.075 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0014$0.001218.0 s
Summarise a 30-page report12,000 / 600$0.011$0.005219.5 s
Code edit6,000 / 1,500$0.010$0.007124.3 s
Agentic coding session60,000 / 4,000$0.060$0.03037.4 s
Structured extraction2,000 / 200$0.0023$0.001217.4 s

See also

Data as of 9 Sept 2026. Compare these configurations.