BenchLeader
GoogleReasoning model

Gemini 3 Flash

Best configuration ranks #66 of 610 on the BenchLeader Index at 61.5 ±3.3 (thinking reasoning effort). Released 17 Dec 2025.

Blended price
$1.13/M
$0.500 in · $3.00 out
Output speed
200 tok/s
First answer
6.12 s
first token 0.85 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index62
  2. Agents & tools67
  3. Knowledge69
  4. Instruction following75
  5. Multimodal65
  6. Long context66
  7. Composite62

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
minimal53.4#226202 tok/s0.85 s$0.0011Coding 51 · Human preference 65 · Multimodal 61 · Reasoning 42
low202 tok/s0.85 s$0.0011Reasoning 35
medium202 tok/s0.85 s$0.0011Reasoning 43
high58.5#121202 tok/s0.85 s$0.0011Agents & tools 52 · Coding 63 · Knowledge 64 · Maths 63 · Reasoning 58
thinkingbest61.5#66200 tok/s6.12 s$0.0011Agents & tools 67 · Composite 62 · Instruction following 75 · Knowledge 69 · Long context 66 · Multimodal 65
default56.6#153202 tok/s0.85 s$0.0011Agents & tools 53 · Coding 53 · Composite 52 · Human preference 67 · Instruction following 55 · Knowledge 62 · Long context 54 · Maths 53 · Multimodal 64 · Reasoning 61

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
GPQA Diamond89.4%#4383.2%#97Epoch AI Benchmarking Hub
SimpleBench61.1%#26SimpleBench
LMArena Hard Prompts1475#621493#35LMArena
GPQA Diamond (AA)not in index89.8%#6881.2%#177Artificial Analysis
Humanity's Last Exam (AA)not in index36.6%#7415.0%#194Artificial Analysis
GPQA Diamond (Vals)not in index87.9%#36Vals AI
ARC-AGI-121.5%#15429.0%#14657.7%#11284.7%#76ARC Prize
ARC-AGI-23.3%#1341.3%#15312.8%#10233.6%#83ARC Prize

Coding

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
SWE-bench Verified (Epoch)75.4%#13Epoch AI Benchmarking Hub
WeirdML61.6%#42WeirdML
GSO-Bench9.8%#17GSO-Bench
LMArena Coding1491#761507#52LMArena
LMArena WebDev1383#801438#58LMArena
LiveCodeBench85.6%#29Vals AI
IOI39.1%#13Vals AI
SWE-bench (Vals)not in index75.0%#42Vals AI
SWE-Bench Pro34.6%#15Scale AI SEAL
SWE-bench Verified (bash only)75.8%#2SWE-bench
SWE-bench Verified (any scaffold)not in index75.8%#5SWE-bench

Agents & tools

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
Terminal-Bench64.3%#11Terminal-Bench
APEX-Agents24.0%#31Mercor
Terminal-Bench Hard38.6%#5931.8%#99Artificial Analysis
τ²-Bench Telecom (AA)not in index80.4%#11943.3%#210Artificial Analysis
Terminal-Bench 2.1 (Vals)53.9%#42Vals AI
MCP Atlas62.0%#23Scale AI SEAL
τ²-bench83.5%#4τ²-bench

Maths

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
FrontierMath Tiers 1–351.2%#38Epoch AI Benchmarking Hub
FrontierMath Tier 417.1%#42Epoch AI Benchmarking Hub
OTIS Mock AIME95.6%#2992.8%#50Epoch AI Benchmarking Hub
ProofBench15.0%#44Vals AI
AIME (Vals)95.6%#8Vals AI
MGSM93.3%#10Vals AI
AIME 202696.7%#6MathArena
HMMT February 202689.4%#14MathArena
MathArena Apex15.6%#16MathArena

Knowledge

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
SimpleQA Verified66.8%#8Epoch AI Benchmarking Hub
AA-Omniscience10.1#64-4.3#113Artificial Analysis
MMLU-Pro88.6%#19Vals AI
LegalBench86.9%#9Vals AI
CorpFin66.4%#19Vals AI
TaxEval73.9%#44Vals AI
MedQA95.8%#11Vals AI

Instruction following

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
IFBench78.0%#1455.1%#145Artificial Analysis

Human preference

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
LMArena Text1459#531474#30LMArena

Multimodal

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
LMArena Vision1266#371285#24LMArena
MMMU-Pro79.9%#3778.5%#48Artificial Analysis

Long context

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
AA-LCR78.0%#8255.3%#233Artificial Analysis

Composite

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
Epoch Capabilities Indexnot in index151.8#38Epoch AI Benchmarking Hub
AA Intelligence Index26.3#10817.9#195Artificial Analysis
Vals Indexnot in index29.4#42Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Google AI Studio Priority84 tok/s1.25 s$0.900$5.401.0M
Google AI Studio83 tok/s1.34 s$0.500$3.001.0M
Google Vertex Priority78 tok/s1.63 s$0.900$5.401.0M
Google Vertex58 tok/s1.19 s$0.500$3.001.0M
Google Vertex Flex50 tok/s19 s$0.250$1.501.0M
Google AI Studio Flex7 tok/s1.94 s$0.250$1.501.0M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.050 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0011$0.00107.6 s
Summarise a 30-page report12,000 / 600$0.0078$0.00389.1 s
Code edit6,000 / 1,500$0.0075$0.005513.6 s
Agentic coding session60,000 / 4,000$0.042$0.02226.1 s
Structured extraction2,000 / 200$0.0016$0.00097.1 s

See also

Data as of 9 Sept 2026. Compare these configurations.