BenchLeader
OpenAIAuto-detected

GPT 4 0613

GPT 4 0613 is an OpenAI proprietary model, released 13 Jun 2023. Its best configuration ranks #549 of 372 on the BenchLeader Index at 39.2 ±6.4, in the lower half. It scores highest in human preference (43) and lowest in maths (30). At $37.50 per million tokens blended it is among the most expensive ranked models. Last measured 2 Sept 2026.

Blended price
$37.50/M
$30.00 in · $60.00 out
Output speed
First answer
Context
8k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index39
  2. Reasoning32
  3. Coding34
  4. Maths30
  5. Human preference43

Versions

OpenAI has shipped 14 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
GPT-5.69 Jul 2026
GPT-5.523 Apr 202666.7#17
GPT-5.5xhigh23 Apr 2026
GPT-5.45 Mar 202663.5#46
GPT-5.112 Nov 202559.6#100
GPT-57 Aug 202559.9#97
GPT-4.114 Apr 202547.4#358
GPT-4.527 Feb 202549.8#310
GPT 4 012525 Jan 202445.5#396
GPT-46 Nov 202344.5#424
GPT 4 0613this page13 Jun 202339.2#549
GPT 4 031414 Mar 202342.0#485
GPT-5.261.6#65
GPT-3.5

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Coding

BenchmarkdefaultSource
WeirdML12.4%#147WeirdML
LMArena Coding1314#250LMArena

Human preference

BenchmarkdefaultSource
LMArena Text1276#260LMArena

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.030
Summarise a 30-page report12,000 / 600$0.396
Code edit6,000 / 1,500$0.270
Agentic coding session60,000 / 4,000$2.04
Structured extraction2,000 / 200$0.072

See also

Data as of 10 Sept 2026. Compare with another model.