BenchLeader
GoogleReasoning modelAuto-detected

Gemini 3.7 Flash

Best configuration ranks #34 of 610 on the BenchLeader Index at 64.4 ±3.9 (medium reasoning effort). Last measured 13 Aug 2026. Released 13 Aug 2026.

Blended price
$1.50/M
$0.750 in · $3.75 out
Output speed
282 tok/s
First answer
5.27 s
first token 1.55 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index64
  2. Reasoning64
  3. Coding68
  4. Knowledge75
  5. Multimodal69
  6. Long context68
  7. Composite79

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
low62.5#53303 tok/s0.62 s$0.0014Coding 63 · Composite 76 · Knowledge 75 · Long context 66 · Multimodal 70 · Reasoning 59
mediumbest64.4#34282 tok/s5.27 s$0.0014Coding 68 · Composite 79 · Knowledge 75 · Long context 68 · Multimodal 69 · Reasoning 64
high60.8#75297 tok/s9.42 s$0.0014Agents & tools 40 · Coding 66 · Composite 63 · Human preference 69 · Knowledge 67 · Maths 63 · Reasoning 69
default61.7#63297 tok/s9.42 s$0.0014Agents & tools 52 · Coding 60 · Composite 79 · Knowledge 77 · Long context 67 · Maths 62 · Multimodal 70

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarklowmediumhighdefaultSourceTrend
GPQA Diamond94.8%#3Epoch AI Benchmarking Hub
LMArena Hard Prompts1508#14LMArena
LiveBench Reasoningnot in index87.8%#18LiveBench
GPQA Diamond (AA)not in index90.1%#6492.1%#3994.5%#6Artificial Analysis
Humanity's Last Exam (AA)not in index35.1%#8039.0%#6147.9%#19Artificial Analysis
GPQA Diamond (Vals)not in index93.9%#5Vals AI
ARC-AGI-185.2%#7591.2%#4595.5%#20ARC Prize
ARC-AGI-252.9%#6963.8%#5084.6%#19ARC Prize

Coding

BenchmarklowmediumhighdefaultSourceTrend
SciCode53.6%#3757.9%#856.8%#11SciCode
FrontierCode43.6%#9Cognition
LMArena Coding1522#23LMArena
LMArena WebDev1587#17LMArena
LiveBench Codingnot in index78.9%#19LiveBench
SciCode (AA)not in index55.7%#2859.8%#457.2%#15Artificial Analysis
LiveCodeBench88.7%#5Vals AI
SWE-bench (Vals)not in index80.8%#23Vals AI

Agents & tools

BenchmarklowmediumhighdefaultSourceTrend
Terminal-Bench11.2%#64Terminal-Bench
Remote Labor Index5.0%#4Scale AI / CAIS
LiveBench Agentic Codingnot in index58.3%#15LiveBench
Terminal-Bench 2.1 (Vals)77.5%#11Vals AI

Maths

BenchmarklowmediumhighdefaultSourceTrend
FrontierMath Tiers 1–371.6%#16Epoch AI Benchmarking Hub
FrontierMath Tier 436.6%#23Epoch AI Benchmarking Hub
OTIS Mock AIME97.2%#23Epoch AI Benchmarking Hub
ProofBench58.0%#11Vals AI
LiveBench Mathematicsnot in index93.5%#12LiveBench

Knowledge

BenchmarklowmediumhighdefaultSourceTrend
SimpleQA Verified69.2%#7Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index68.0%#42LiveBench
AA-Omniscience22.1#3423.7#3226.5#28Artificial Analysis
MMLU-Pro90.1%#6Vals AI
LegalBench87.3%#4Vals AI
TaxEval74.7%#30Vals AI

Instruction following

BenchmarklowmediumhighdefaultSourceTrend
LiveBench Languagenot in index85.5%#9LiveBench

Human preference

BenchmarklowmediumhighdefaultSourceTrend
LMArena Text1491#11LMArena

Multimodal

BenchmarklowmediumhighdefaultSourceTrend
MMMU-Pro84.9%#784.7%#885.5%#5Artificial Analysis

Long context

BenchmarklowmediumhighdefaultSourceTrend
AA-LCR78.7%#7583.0%#1281.7%#30Artificial Analysis

Composite

BenchmarklowmediumhighdefaultSourceTrend
Epoch Capabilities Indexnot in index157.4#12Epoch AI Benchmarking Hub
LiveBench78.8%#9LiveBench
AA Intelligence Index37.0#4939.6#3839.4#40Artificial Analysis
Vals Indexnot in index59.3#13Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Google AI Studio Priority190 tok/s1.33 s$1.35$6.751.0M
Google AI Studio Flex175 tok/s1.12 s$0.375$1.881.0M
Google AI Studio162 tok/s1.34 s$0.750$3.751.0M
Google Vertex83 tok/s2.27 s$0.750$3.751.0M
Google Vertex Priority79 tok/s1.75 s$1.35$6.751.0M
Google Vertex Flex27 tok/s13 s$0.375$1.881.0M

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$1.03$2.06$3.09$4.13Aug 26
input outputnow $0.750 in · $3.75 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.075 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0014$0.00126.3 s
Summarise a 30-page report12,000 / 600$0.011$0.00527.4 s
Code edit6,000 / 1,500$0.010$0.007110.6 s
Agentic coding session60,000 / 4,000$0.060$0.03019.4 s
Structured extraction2,000 / 200$0.0023$0.00126.0 s

See also

Data as of 9 Sept 2026. Compare these configurations.