BenchLeader
GoogleReasoning modelAuto-detected

Gemini 3.5 Flash

Best configuration ranks #39 of 610 on the BenchLeader Index at 63.9 ±2.7 (medium reasoning effort). Last measured 8 Sept 2026. Released 19 May 2026.

Blended price
$3.38/M
$1.50 in · $9.00 out
Output speed
209 tok/s
First answer
17 s
first token 1.66 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index64
  2. Reasoning67
  3. Coding59
  4. Agents & tools68
  5. Knowledge74
  6. Instruction following72
  7. Human preference67
  8. Multimodal68
  9. Long context64
  10. Composite71

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
minimal57.3#142187 tok/s1.02 s$0.0033Agents & tools 74 · Composite 59 · Instruction following 49 · Knowledge 64 · Long context 57 · Maths 59 · Multimodal 65 · Reasoning 49
low209 tok/s17 s$0.0033Maths 63 · Reasoning 64
mediumbest63.9#39209 tok/s17 s$0.0033Agents & tools 68 · Coding 59 · Composite 71 · Human preference 67 · Instruction following 72 · Knowledge 74 · Long context 64 · Multimodal 68 · Reasoning 67
high60.4#86209 tok/s17 s$0.0033Agents & tools 64 · Coding 61 · Composite 50 · Human preference 68 · Knowledge 64 · Maths 57 · Multimodal 67 · Reasoning 62
default60.7#78209 tok/s17 s$0.0033Agents & tools 57 · Composite 71 · Human preference 33 · Instruction following 74 · Knowledge 74 · Long context 63 · Maths 58 · Multimodal 69 · Reasoning 72

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkminimallowmediumhighdefaultSourceTrend
GPQA Diamond86.4%#6688.9%#4592.8%#19Epoch AI Benchmarking Hub
SimpleBench76.7%#6SimpleBench
LMArena Hard Prompts1493#361498#29LMArena
LiveBench Reasoningnot in index82.0%#36LiveBench
GPQA Diamond (AA)not in index82.8%#16392.1%#3992.2%#38Artificial Analysis
Humanity's Last Exam (AA)not in index24.1%#13841.3%#4742.7%#36Artificial Analysis
GPQA Diamond (Vals)not in index92.7%#14Vals AI
EnigmaEval25.4%#4Scale AI SEAL
ARC-AGI-148.8%#12392.5%#35ARC Prize
ARC-AGI-28.9%#10872.1%#34ARC Prize

Coding

BenchmarkminimallowmediumhighdefaultSourceTrend
SWE-bench Verified (Epoch)79.3%#3Epoch AI Benchmarking Hub
SciCode53.1%#45SciCode
WeirdML62.6%#41WeirdML
LMArena Coding1503#561508#51LMArena
LMArena WebDev1491#421500#40LMArena
LiveBench Codingnot in index78.2%#22LiveBench
SciCode (AA)not in index53.9%#45Artificial Analysis
LiveCodeBench87.6%#12Vals AI
SWE-bench (Vals)not in index78.8%#28Vals AI

Agents & tools

BenchmarkminimallowmediumhighdefaultSourceTrend
LiveBench Agentic Codingnot in index49.0%#33LiveBench
Terminal-Bench Hard46.2%#3039.4%#5640.9%#51Artificial Analysis
τ²-Bench Telecom (AA)not in index58.8%#17795.6%#2095.3%#24Artificial Analysis
Terminal-Bench 2.1 (Vals)74.2%#15Vals AI
MCP Atlas83.6%#4Scale AI SEAL
HiL-Bench27.7%#12Scale AI SEAL

Maths

BenchmarkminimallowmediumhighdefaultSourceTrend
FrontierMath Tiers 1–362.8%#26Epoch AI Benchmarking Hub
FrontierMath Tier 426.8%#31Epoch AI Benchmarking Hub
OTIS Mock AIME80.0%#9988.9%#6195.6%#29Epoch AI Benchmarking Hub
ProofBench31.0%#28Vals AI
LiveBench Mathematicsnot in index88.2%#30LiveBench
AIME 202695.0%#18MathArena
HMMT February 202695.5%#6MathArena
MathArena Apex32.3%#9MathArena

Knowledge

BenchmarkminimallowmediumhighdefaultSourceTrend
SimpleQA Verified66.2%#9Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index64.9%#46LiveBench
AA-Omniscience0.7#9120.8#4021.2#38Artificial Analysis
MMLU-Pro89.5%#10Vals AI
LegalBench83.6%#47Vals AI
CorpFin64.7%#35Vals AI
TaxEval74.4%#36Vals AI

Instruction following

BenchmarkminimallowmediumhighdefaultSourceTrend
LiveBench Languagenot in index84.6%#11LiveBench
IFBench47.3%#18674.6%#3576.3%#21Artificial Analysis

Human preference

BenchmarkminimallowmediumhighdefaultSourceTrend
LMArena Text1476#271479#23LMArena
EQ-Bench 41087#23EQ-Bench

Multimodal

BenchmarkminimallowmediumhighdefaultSourceTrend
LMArena Vision1306#91311#8LMArena
MMMU-Pro80.1%#3683.9%#1584.3%#12Artificial Analysis

Long context

BenchmarkminimallowmediumhighdefaultSourceTrend
AA-LCR61.3%#21074.3%#11673.3%#129Artificial Analysis

Composite

BenchmarkminimallowmediumhighdefaultSourceTrend
Epoch Capabilities Indexnot in index154.7#26Epoch AI Benchmarking Hub
LiveBench74.6%#29LiveBench
AA Intelligence Index23.9#12633.6#6533.0#67Artificial Analysis
Vals Indexnot in index53.1#24Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Google AI Studio Priority192 tok/s2.13 s$2.70$16.201.0M
Google AI Studio187 tok/s1.75 s$1.50$9.001.0M
Google AI Studio Flex158 tok/s1.41 s$0.750$4.501.0M
Google Vertex127 tok/s1.57 s$1.50$9.001.0M
Google Vertex Priority36 tok/s1.48 s$2.70$16.201.0M
Google Vertex Flex4 tok/s13 s$0.750$4.501.0M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.150 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0033$0.002918.4 s
Summarise a 30-page report12,000 / 600$0.023$0.01119.8 s
Code edit6,000 / 1,500$0.022$0.01624.1 s
Agentic coding session60,000 / 4,000$0.126$0.06536.1 s
Structured extraction2,000 / 200$0.0048$0.002817.9 s

See also

Data as of 9 Sept 2026. Compare these configurations.