BenchLeader
OpenAIAuto-detected

GPT 3.5 Turbo 0125

GPT 3.5 Turbo 0125 is an OpenAI proprietary model, released 25 Jan 2024. Its best configuration ranks #607 of 372 on the BenchLeader Index at 35.5 ±4.9, in the lower half. It scores highest in human preference (37) and lowest in maths (27). At $0.750 per million tokens blended it is mid-priced. Last measured 2 Sept 2026.

Blended price
$0.750/M
$0.500 in · $1.50 out
Output speed
First answer
Context
16k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index36
  2. Reasoning28
  3. Coding28
  4. Maths27
  5. Human preference37

Versions

OpenAI has shipped 4 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
GPT-4 Turbo9 Apr 202441.1#508
GPT 3.5 Turbo 0125this page25 Jan 202435.5#607
GPT 3.5 Turbo 11066 Nov 202339.3#545
GPT-3.5 Turbohigh21 Sept 2023

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Coding

BenchmarkdefaultSource
WeirdML3.5%#154WeirdML
LMArena Coding1275#277LMArena

Human preference

BenchmarkdefaultSource
LMArena Text1225#286LMArena

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0007
Summarise a 30-page report12,000 / 600$0.0069
Code edit6,000 / 1,500$0.0053
Agentic coding session60,000 / 4,000$0.036
Structured extraction2,000 / 200$0.0013

See also

Data as of 10 Sept 2026. Compare with another model.