BenchLeader
OpenAIReasoning model

o1

Best configuration ranks #230 of 610 on the BenchLeader Index at 53.2 ±4.3. Last measured 2 Sept 2026. Released 17 Dec 2024.

Blended price
$26.25/M
$15.00 in · $60.00 out
Output speed
First answer
Context
200k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index53
  2. Reasoning43
  3. Coding56
  4. Agents & tools42
  5. Maths50
  6. Knowledge59
  7. Instruction following69
  8. Human preference59
  9. Multimodal55
  10. Long context59
  11. Composite48

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
low$0.024Maths 37 · Reasoning 55
medium48.3#330$0.024Long context 48 · Maths 47 · Reasoning 48
high44.3#417$0.024Agents & tools 24 · Coding 38 · Knowledge 55 · Maths 50 · Reasoning 49
defaultbest53.2#230$0.024Agents & tools 42 · Coding 56 · Composite 48 · Human preference 59 · Instruction following 69 · Knowledge 59 · Long context 59 · Maths 50 · Multimodal 55 · Reasoning 43

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarklowmediumhighdefaultSource
GPQA Diamond74.2%#14175.8%#13176.8%#12850.3%#200Epoch AI Benchmarking Hub
Humanity's Last Exam8.0%#32Scale AI / CAIS
SimpleBench36.7%#6240.1%#6041.7%#56SimpleBench
LMArena Hard Prompts1418#134LMArena
GPQA Diamond (AA)not in index76.5%#225Artificial Analysis
Humanity's Last Exam (AA)not in index7.0%#303Artificial Analysis
GPQA Diamond (Vals)not in index73.2%#83Vals AI

Coding

BenchmarklowmediumhighdefaultSource
WeirdML46.1%#7847.6%#73WeirdML
LMArena Coding1433#147LMArena
LiveCodeBench50.3%#111Vals AI
Aider Polyglot61.7%#12Aider polyglot leaderboard
SWE-bench Verified (any scaffold)not in index64.6%#23SWE-bench

Agents & tools

BenchmarklowmediumhighdefaultSource
Cybench10.0%#15Cybench
APEX-Agents1.1%#62Mercor
Terminal-Bench Hard12.9%#199Artificial Analysis
τ²-Bench Telecom (AA)not in index62.6%#171Artificial Analysis

Maths

BenchmarklowmediumhighdefaultSource
FrontierMath Tiers 1–38.4%#9210.2%#9014.7%#87Epoch AI Benchmarking Hub
OTIS Mock AIME53.3%#16173.3%#11473.3%#11431.1%#187Epoch AI Benchmarking Hub
MATH Level 594.4%#1794.7%#1681.7%#32Epoch AI Benchmarking Hub
AIME (Vals)71.5%#53Vals AI
MGSM89.3%#45Vals AI

Knowledge

BenchmarklowmediumhighdefaultSource
SimpleQA Verified41.1%#35Epoch AI Benchmarking Hub
AA-Omniscience-11.1#150Artificial Analysis
MMLU-Pro83.5%#71Vals AI
LegalBench80.4%#79Vals AI
TaxEval74.3%#38Vals AI
MedQA96.5%#1Vals AI

Instruction following

BenchmarklowmediumhighdefaultSource
IFBench70.3%#68Artificial Analysis

Human preference

BenchmarklowmediumhighdefaultSource
LMArena Text1402#130LMArena

Multimodal

BenchmarklowmediumhighdefaultSource
LMArena Vision1168#90LMArena
VISTA45.3%#28Scale AI SEAL
MMMU (validation)78.2%#9MMMU

Long context

BenchmarklowmediumhighdefaultSource
Fiction.LiveBench 120k53.1%#20Fiction.live
AA-LCR65.0%#193Artificial Analysis

Composite

BenchmarklowmediumhighdefaultSource
Epoch Capabilities Indexnot in index141.9#89Epoch AI Benchmarking Hub
AA Intelligence Index15.2#220Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $7.50 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.024$0.022
Summarise a 30-page report12,000 / 600$0.216$0.148
Code edit6,000 / 1,500$0.180$0.146
Agentic coding session60,000 / 4,000$1.14$0.802
Structured extraction2,000 / 200$0.042$0.031

See also

Data as of 9 Sept 2026. Compare these configurations.