BenchLeader
GoogleReasoning model

Gemini 3 Pro

Best configuration ranks #64 of 610 on the BenchLeader Index at 61.6 ±3.1. Last measured 8 Sept 2026. Released 18 Nov 2025.

Blended price
$4.50/M
$2.00 in · $12.00 out
Output speed
First answer
Context
1.0M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index62
  2. Reasoning67
  3. Coding59
  4. Agents & tools62
  5. Maths48
  6. Knowledge61
  7. Instruction following64
  8. Human preference69
  9. Multimodal66
  10. Long context65
  11. Composite64

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
low56.2#161$0.0044Agents & tools 63 · Composite 57 · Instruction following 51 · Knowledge 65 · Long context 63
high57.0#151$0.0044Agents & tools 54 · Coding 61 · Knowledge 60 · Maths 63
defaultbest61.6#64$0.0044Agents & tools 62 · Coding 59 · Composite 64 · Human preference 69 · Instruction following 64 · Knowledge 61 · Long context 65 · Maths 48 · Multimodal 66 · Reasoning 67

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarklowhighdefaultSourceTrend
GPQA Diamond92.6%#21Epoch AI Benchmarking Hub
Humanity's Last Exam37.5%#5Scale AI / CAIS
SimpleBench76.4%#7SimpleBench
LMArena Hard Prompts1504#20LMArena
GPQA Diamond (AA)not in index88.7%#8690.8%#55Artificial Analysis
Humanity's Last Exam (AA)not in index29.5%#10839.7%#57Artificial Analysis
GPQA Diamond (Vals)not in index91.7%#18Vals AI
Kagi LLM Benchmark80.1%#9Kagi LLM Benchmark
ARC-AGI-175.0%#92ARC Prize
ARC-AGI-254.0%#68ARC Prize

Coding

BenchmarklowhighdefaultSourceTrend
SWE-bench Verified (Epoch)72.9%#21Epoch AI Benchmarking Hub
WeirdML69.9%#31WeirdML
GSO-Bench18.6%#13GSO-Bench
LMArena Coding1518#32LMArena
LMArena WebDev1439#57LMArena
LiveCodeBench86.4%#22Vals AI
IOI38.8%#14Vals AI
SWE-bench (Vals)not in index76.4%#35Vals AI
SWE-Bench Pro43.3%#8Scale AI SEAL
SWE-bench Verified (bash only)69.6%#1374.2%#6SWE-bench
SWE-bench Verified (any scaffold)not in index77.4%#3SWE-bench

Agents & tools

BenchmarklowhighdefaultSourceTrend
Terminal-Bench69.4%#7Terminal-Bench
GDPval40.3%#5OpenAI
Remote Labor Index1.3%#11Scale AI / CAIS
APEX-Agents31.5%#27Mercor
Terminal-Bench Hard34.1%#8741.7%#48Artificial Analysis
τ²-Bench Telecom (AA)not in index68.1%#15787.1%#79Artificial Analysis
MCP Atlas70.3%#18Scale AI SEAL
BFCL Overall72.5%#3Berkeley Function Calling Leaderboard
τ²-bench82.5%#5τ²-bench

Maths

BenchmarklowhighdefaultSourceTrend
OTIS Mock AIME91.4%#53Epoch AI Benchmarking Hub
ProofBench20.0%#36Vals AI
AIME (Vals)96.7%#4Vals AI
MGSM93.9%#7Vals AI
AIME 202691.7%#24MathArena
HMMT February 202686.4%#20MathArena
MathArena Apex23.4%#14MathArena

Knowledge

BenchmarklowhighdefaultSourceTrend
AA-Omniscience1.8#8315.3#51Artificial Analysis
MMLU-Pro90.1%#7Vals AI
LegalBench87.0%#5Vals AI
CorpFin63.7%#42Vals AI
TaxEval72.6%#61Vals AI
MedQA96.0%#8Vals AI
PRBench Finance39.2%#24Scale AI SEAL
PRBench Legal40.6%#23Scale AI SEAL
MultiNRC59.0%#6Scale AI SEAL

Instruction following

BenchmarklowhighdefaultSourceTrend
IFBench49.7%#17470.4%#67Artificial Analysis
MultiChallenge65.7%#5Scale AI SEAL
TutorBench53.7%#9Scale AI SEAL

Human preference

BenchmarklowhighdefaultSourceTrend
LMArena Text1486#16LMArena

Multimodal

BenchmarklowhighdefaultSourceTrend
LMArena Vision1305#11LMArena
MMMU-Pro80.2%#35Artificial Analysis
VISTA51.5%#7Scale AI SEAL
MMMU-Pro (official)not in index81.0%#1MMMU

Long context

BenchmarklowhighdefaultSourceTrend
AA-LCR74.0%#12076.0%#102Artificial Analysis

Composite

BenchmarklowhighdefaultSourceTrend
Epoch Capabilities Indexnot in index153.0#33Epoch AI Benchmarking Hub
AA Intelligence Index22.3#14728.0#94Artificial Analysis

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0044
Summarise a 30-page report12,000 / 600$0.031
Code edit6,000 / 1,500$0.030
Agentic coding session60,000 / 4,000$0.168
Structured extraction2,000 / 200$0.0064

See also

Data as of 9 Sept 2026. Compare these configurations.