BenchLeader
OpenAIAuto-detected

GPT 3.5 Turbo 1106

GPT 3.5 Turbo 1106 is an OpenAI proprietary model, released 6 Nov 2023. Its best configuration ranks #545 of 372 on the BenchLeader Index at 39.3 ±4.3, in the lower half. It scores highest in coding (38) and lowest in reasoning (28). At $1.25 per million tokens blended it is pricier than most ranked models. Last measured 2 Sept 2026.

Blended price
$1.25/M
$1.00 in · $2.00 out
Output speed
First answer
Context
16k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index39
  2. Reasoning28
  3. Coding38
  4. Maths31
  5. Human preference35

Versions

OpenAI has shipped 4 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
GPT-4 Turbo9 Apr 202441.1#508
GPT 3.5 Turbo 012525 Jan 202435.5#607
GPT 3.5 Turbo 1106this page6 Nov 202339.3#545
GPT-3.5 Turbohigh21 Sept 2023

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Coding

BenchmarkdefaultSource
LMArena Coding1262#286LMArena

Maths

BenchmarkdefaultSource
MATH Level 515.9%#81Epoch AI Benchmarking Hub

Human preference

BenchmarkdefaultSource
LMArena Text1203#296LMArena

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0010
Summarise a 30-page report12,000 / 600$0.013
Code edit6,000 / 1,500$0.0090
Agentic coding session60,000 / 4,000$0.068
Structured extraction2,000 / 200$0.0024

See also

Data as of 10 Sept 2026. Compare with another model.