BenchLeader
OpenAIReasoning model

GPT-5

Best configuration ranks #100 of 610 on the BenchLeader Index at 59.5 ±3.5. Last measured 2 Sept 2026. Released 7 Aug 2025.

Blended price
$3.44/M
$1.25 in · $10.00 out
Output speed
76 tok/s
First answer
74 s
first token 6.54 s
Context
400k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index60
  2. Reasoning61
  3. Coding61
  4. Agents & tools51
  5. Knowledge62
  6. Instruction following68
  7. Human preference62
  8. Multimodal60
  9. Long context66
  10. Composite58

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
minimal46.0#37479 tok/s1.35 s$0.0035Agents & tools 49 · Composite 43 · Instruction following 47 · Knowledge 48 · Maths 40 · Multimodal 46 · Reasoning 40
low53.1#23372 tok/s9.74 s$0.0035Agents & tools 57 · Composite 55 · Instruction following 65 · Knowledge 59 · Maths 48 · Multimodal 58 · Reasoning 37
medium58.5#12281 tok/s42 s$0.0035Agents & tools 57 · Coding 52 · Composite 58 · Instruction following 69 · Knowledge 59 · Long context 71 · Maths 66 · Multimodal 59 · Reasoning 49
high54.5#19576 tok/s74 s$0.0035Agents & tools 43 · Coding 55 · Human preference 62 · Knowledge 59 · Maths 57 · Multimodal 54 · Reasoning 55
thinking76 tok/s74 s$0.0035Instruction following 60
defaultbest59.5#10076 tok/s74 s$0.0035Agents & tools 51 · Coding 61 · Composite 58 · Human preference 62 · Instruction following 68 · Knowledge 62 · Long context 66 · Multimodal 60 · Reasoning 61

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
GPQA Diamond71.7%#15085.3%#7686.2%#70Epoch AI Benchmarking Hub
Humanity's Last Exam25.3%#1125.3%#11Scale AI / CAIS
SimpleBench56.7%#36SimpleBench
LMArena Hard Prompts1448#951449#93LMArena
GPQA Diamond (AA)not in index67.3%#30680.8%#18484.2%#14785.3%#127Artificial Analysis
Humanity's Last Exam (AA)not in index6.0%#33319.6%#16325.4%#13328.5%#115Artificial Analysis
GPQA Diamond (Vals)not in index85.6%#44Vals AI
Kagi LLM Benchmark72.7%#21Kagi LLM Benchmark
ARC-AGI-16.0%#17244.0%#12856.2%#11665.7%#101ARC Prize
ARC-AGI-20.0%#1741.9%#1447.5%#1119.9%#106ARC Prize

Coding

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
SWE-bench Verified (Epoch)71.5%#2373.5%#19Epoch AI Benchmarking Hub
SciCode42.9%#94SciCode
WeirdML60.7%#4539.8%#102WeirdML
GSO-Bench6.9%#20GSO-Bench
LMArena Coding1470#1011463#110LMArena
LMArena WebDev1420#64LMArena
LiveCodeBench85.9%#27Vals AI
IOI20.0%#24Vals AI
SWE-bench (Vals)not in index69.0%#64Vals AI
SWE-Bench Pro41.8%#10Scale AI SEAL
Aider Polyglot88.0%#1Aider polyglot leaderboard
SWE-bench Verified (bash only)65.0%#19SWE-bench
SWE-bench Verified (any scaffold)not in index75.6%#7SWE-bench

Agents & tools

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
Terminal-Bench49.6%#2649.6%#26Terminal-Bench
GDPval34.8%#6OpenAI
Remote Labor Index1.7%#10Scale AI / CAIS
APEX-Agents18.3%#3918.3%#39Mercor
Terminal-Bench Hard18.2%#16626.5%#12637.9%#6132.6%#95Artificial Analysis
τ²-Bench Telecom (AA)not in index67.0%#16084.2%#10186.5%#8584.8%#95Artificial Analysis

Maths

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
FrontierMath Tiers 1–318.3%#8037.2%#4955.4%#35Epoch AI Benchmarking Hub
FrontierMath Tier 421.9%#38Epoch AI Benchmarking Hub
OTIS Mock AIME46.7%#16987.2%#6991.4%#53Epoch AI Benchmarking Hub
MATH Level 597.9%#298.1%#1Epoch AI Benchmarking Hub
ProofBench18.0%#38Vals AI
AIME (Vals)93.4%#14Vals AI
MGSM92.8%#14Vals AI
IMO 202538.1%#1MathArena
MathArena Apex1.0%#35MathArena

Knowledge

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
SimpleQA Verified50.1%#21Epoch AI Benchmarking Hub
AA-Omniscience-33.8#229-10.8#146-10.9#147-8.7#130Artificial Analysis
MMLU-Pro86.5%#40Vals AI
LegalBench86.0%#14Vals AI
CorpFin61.1%#59Vals AI
TaxEval73.4%#48Vals AI
MedQA96.3%#4Vals AI
PRBench Finance51.3%#5Scale AI SEAL
PRBench Legal49.0%#10Scale AI SEAL
MultiNRC52.1%#9Scale AI SEAL

Instruction following

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
IFBench45.6%#19966.6%#9270.6%#6673.1%#44Artificial Analysis
MultiChallenge63.2%#7Scale AI SEAL
TutorBench55.3%#4Scale AI SEAL

Human preference

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
LMArena Text1434#881427#97LMArena

Multimodal

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
LMArena Vision1209#701232#61LMArena
MMMU-Pro62.1%#16673.8%#9374.3%#8474.2%#86Artificial Analysis
VISTA49.7%#11Scale AI SEAL

Long context

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
Fiction.LiveBench 120k96.9%#2Fiction.live
AA-LCR76.0%#10278.2%#81Artificial Analysis

Composite

BenchmarkminimallowmediumhighthinkingdefaultSourceTrend
Epoch Capabilities Indexnot in index150#45Epoch AI Benchmarking Hub
AA Intelligence Index11.4#28620.8#16722.9#13823.0#136Artificial Analysis

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
OpenAI62 tok/s3.81 s$1.25$10.00400k
Azure48 tok/s9.27 s$1.25$10.00400k

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$3.03$6.05$9.08$12.10Jan 26Feb 26Mar 26Apr 26May 26Jun 26Jul 26Aug 26
input outputnow $1.38 in · $11.00 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.125 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0035$0.00321.3 min
Summarise a 30-page report12,000 / 600$0.021$0.0111.4 min
Code edit6,000 / 1,500$0.022$0.0171.6 min
Agentic coding session60,000 / 4,000$0.115$0.0642.1 min
Structured extraction2,000 / 200$0.0045$0.00281.3 min

See also

Data as of 9 Sept 2026. Compare these configurations.