BenchLeader

GPT-4o

Best configuration ranks #461 of 610 on the BenchLeader Index at 42.3 ±5.4. Last measured 2 Sept 2026. Released 13 May 2024.

Blended price
$4.38/M
$2.50 in · $10.00 out
Output speed
140 tok/s
First answer
1.04 s
first token 0.77 s
Context
128k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index42
  2. Reasoning37
  3. Coding35
  4. Agents & tools31
  5. Maths32
  6. Knowledge45
  7. Instruction following29
  8. Human preference64
  9. Multimodal49
  10. Long context54
  11. Composite40

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
high41.1#493100 tok/s1.20 s$0.0065Coding 28 · Knowledge 46 · Maths 40
defaultbest42.3#461140 tok/s1.04 s$0.0040Agents & tools 31 · Coding 35 · Composite 40 · Human preference 64 · Instruction following 29 · Knowledge 45 · Long context 54 · Maths 32 · Multimodal 49 · Reasoning 37

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Agents & tools

BenchmarkhighdefaultSource
GDPval9.9%#11OpenAI
Cybench12.5%#14Cybench
APEX-Agents1.1%#62Mercor
Terminal-Bench Hard8.3%#226Artificial Analysis
τ²-Bench Telecom (AA)not in index28.9%#258Artificial Analysis

Knowledge

BenchmarkhighdefaultSource
SimpleQA Verified26.0%#59Epoch AI Benchmarking Hub
AA-Omniscience-10.5#142Artificial Analysis
MMLU-Pro74.1%#109Vals AI
LegalBench82.2%#66Vals AI
CorpFin45.9%#109Vals AI
TaxEval74.5%#34Vals AI
MedQA88.2%#54Vals AI
MultiNRC12.4%#39Scale AI SEAL

Instruction following

BenchmarkhighdefaultSource
IFBench36.0%#293Artificial Analysis
TutorBench36.1%#27Scale AI SEAL

Human preference

BenchmarkhighdefaultSource
LMArena Text1443#75LMArena

Multimodal

BenchmarkhighdefaultSource
LMArena Vision1244#55LMArena
MMMU-Pro56.3%#190Artificial Analysis
VISTA38.0%#42Scale AI SEAL
MMMU (validation)70.7%#21MMMU
MMMU-Pro (official)not in index51.9%#15MMMU

Long context

BenchmarkhighdefaultSource
AA-LCR56.0%#231Artificial Analysis

Composite

BenchmarkhighdefaultSource
Epoch Capabilities Indexnot in index129.0#137Epoch AI Benchmarking Hub
AA Intelligence Index9.0#343Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Azure63 tok/s0.88 s$2.50$10.00128k
OpenAI48 tok/s0.67 s$2.50$10.00128k

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $1.25 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0040$0.00363.2 s
Summarise a 30-page report12,000 / 600$0.036$0.0255.3 s
Code edit6,000 / 1,500$0.030$0.02411.8 s
Agentic coding session60,000 / 4,000$0.190$0.13429.7 s
Structured extraction2,000 / 200$0.0070$0.00512.5 s

See also

Data as of 9 Sept 2026. Compare these configurations.