BenchLeader
OpenAIAuto-detected

GPT-3.5-turbo

Best configuration ranks #586 of 610 on the BenchLeader Index at 36.0 ±4.5. Last measured 2 Sept 2026. Released 25 Jan 2024.

Blended price
$0.750/M
$0.500 in · $1.50 out
Output speed
24 tok/s
First answer
0.58 s
Context
16k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index36
  2. Reasoning28
  3. Coding29
  4. Maths27
  5. Human preference37
  6. Composite36

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
high24 tok/s0.58 s$0.0007Knowledge 26
defaultbest36.0#58624 tok/s0.58 s$0.0007Coding 29 · Composite 36 · Human preference 37 · Maths 27 · Reasoning 28

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkhighdefaultSource
GPQA Diamond28.0%#255Epoch AI Benchmarking Hub
LMArena Hard Prompts1228#278LMArena
GPQA Diamond (AA)not in index29.7%#491Artificial Analysis
GPQA Diamond (Vals)not in index30.6%#128Vals AI

Coding

BenchmarkhighdefaultSource
WeirdML3.5%#148WeirdML
LMArena Coding1275#267LMArena

Knowledge

BenchmarkhighdefaultSource
LegalBench64.4%#124Vals AI
MedQA58.5%#83Vals AI

Human preference

BenchmarkhighdefaultSource
LMArena Text1225#276LMArena

Composite

BenchmarkhighdefaultSource
Epoch Capabilities Indexnot in index118.5#168Epoch AI Benchmarking Hub
AA Intelligence Index5.5#488Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
OpenAI24 tok/s0.58 s$0.500$1.5016k

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.000713.1 s
Summarise a 30-page report12,000 / 600$0.006925.6 s
Code edit6,000 / 1,500$0.00531.1 min
Agentic coding session60,000 / 4,000$0.0362.8 min
Structured extraction2,000 / 200$0.00138.9 s

See also

Data as of 9 Sept 2026. Compare these configurations.