BenchLeader
GoogleReasoning modelAuto-detected

Gemini 3.1 Pro

Best configuration ranks #37 of 610 on the BenchLeader Index at 63.9 ±3.9. Last measured 8 Sept 2026. Released 19 Feb 2026.

Blended price
$4.50/M
$2.00 in · $12.00 out
Output speed
110 tok/s
First answer
30 s
first token 5.50 s
Context
1.0M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index64
  2. Reasoning70
  3. Coding60
  4. Agents & tools64
  5. Maths61
  6. Knowledge66
  7. Instruction following70
  8. Human preference60
  9. Multimodal66
  10. Long context68
  11. Composite67

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
high60.7#76110 tok/s30 s$0.0044Agents & tools 58 · Coding 64 · Composite 57 · Knowledge 66 · Maths 66 · Reasoning 65
thinking110 tok/s30 s$0.0044Coding 60
defaultbest63.9#37110 tok/s30 s$0.0044Agents & tools 64 · Coding 60 · Composite 67 · Human preference 60 · Instruction following 70 · Knowledge 66 · Long context 68 · Maths 61 · Multimodal 66 · Reasoning 70

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkhighthinkingdefaultSourceTrend
GPQA Diamond94.4%#594.1%#7Epoch AI Benchmarking Hub
Humanity's Last Exam46.4%#2Scale AI / CAIS
SimpleBench79.6%#3SimpleBench
LMArena Hard Prompts1507#16LMArena
LiveBench Reasoningnot in index84.0%#31LiveBench
GPQA Diamond (AA)not in index94.1%#7Artificial Analysis
Humanity's Last Exam (AA)not in index47.0%#23Artificial Analysis
GPQA Diamond (Vals)not in index95.5%#1Vals AI
EnigmaEval36.8%#3Scale AI SEAL
ARC-AGI-198.0%#5ARC Prize
ARC-AGI-277.1%#28ARC Prize
ARC-AGI-30.4%#24ARC Prize

Coding

BenchmarkhighthinkingdefaultSourceTrend
SciCode58.9%#4SciCode
WeirdML72.1%#28WeirdML
GSO-Bench22.6%#12GSO-Bench
LMArena Coding1521#25LMArena
LMArena WebDev1447#54LMArena
LiveBench Codingnot in index76.5%#32LiveBench
SciCode (AA)not in index58.7%#10Artificial Analysis
LiveCodeBench88.5%#6Vals AI
SWE-bench (Vals)not in index78.8%#28Vals AI
SWE-Bench Pro46.1%#5Scale AI SEAL

Agents & tools

BenchmarkhighthinkingdefaultSourceTrend
Terminal-Bench80.2%#3Terminal-Bench
APEX-Agents33.5%#22Mercor
LMArena Agent-5.3#33LMArena
LiveBench Agentic Codingnot in index44.1%#42LiveBench
Terminal-Bench Hard53.8%#16Artificial Analysis
τ²-Bench Telecom (AA)not in index95.6%#20Artificial Analysis
Terminal-Bench 2.1 (Vals)70.8%#20Vals AI
MCP Atlas78.2%#12Scale AI SEAL
HiL-Bench35.3%#9Scale AI SEAL

Maths

BenchmarkhighthinkingdefaultSourceTrend
FrontierMath Tiers 1–359.6%#27Epoch AI Benchmarking Hub
FrontierMath Tier 426.8%#31Epoch AI Benchmarking Hub
OTIS Mock AIME95.6%#2995.6%#28Epoch AI Benchmarking Hub
ProofBench26.0%#30Vals AI
LiveBench Mathematicsnot in index91.0%#20LiveBench
AIME (Vals)98.1%#1Vals AI
AIME 202698.3%#4MathArena
HMMT February 202694.7%#8MathArena
MathArena Apex60.9%#5MathArena

Knowledge

BenchmarkhighthinkingdefaultSourceTrend
SimpleQA Verified73.5%#2Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index78.5%#15LiveBench
AA-Omniscience31.9#15Artificial Analysis
MMLU-Pro91.0%#4Vals AI
LegalBench87.4%#3Vals AI
CorpFin64.5%#36Vals AI
TaxEval72.9%#57Vals AI
MedQA96.4%#3Vals AI
PRBench Finance41.9%#21Scale AI SEAL
PRBench Legal44.0%#17Scale AI SEAL
MultiNRC64.7%#3Scale AI SEAL

Instruction following

BenchmarkhighthinkingdefaultSourceTrend
LiveBench Languagenot in index85.4%#10LiveBench
IFBench77.1%#18Artificial Analysis
MultiChallenge71.4%#3Scale AI SEAL
TutorBench53.0%#12Scale AI SEAL

Human preference

BenchmarkhighthinkingdefaultSourceTrend
LMArena Text1487#15LMArena
EQ-Bench 41142#20EQ-Bench

Multimodal

BenchmarkhighthinkingdefaultSourceTrend
LMArena Vision1295#17LMArena
MMMU-Pro82.4%#19Artificial Analysis
MMMU-Pro (official)not in index80.5%#2MMMU

Long context

BenchmarkhighthinkingdefaultSourceTrend
AA-LCR82.0%#25Artificial Analysis

Composite

BenchmarkhighthinkingdefaultSourceTrend
Epoch Capabilities Indexnot in index155#25Epoch AI Benchmarking Hub
LiveBench77.0%#17LiveBench
AA Intelligence Index30.4#80Artificial Analysis
Vals Indexnot in index41.9#34Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Google AI Studio Flex115 tok/s7.05 s$1.00$6.001.0M
Google Vertex102 tok/s2.88 s$2.00$12.001.0M
Google AI Studio90 tok/s5.50 s$2.00$12.001.0M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.200 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0044$0.003933.2 s
Summarise a 30-page report12,000 / 600$0.031$0.01535.9 s
Code edit6,000 / 1,500$0.030$0.02244.1 s
Agentic coding session60,000 / 4,000$0.168$0.0871.1 min
Structured extraction2,000 / 200$0.0064$0.003732.3 s

See also

Data as of 9 Sept 2026. Compare these configurations.