GPT 4 0314
GPT 4 0314 is an OpenAI proprietary model, released 14 Mar 2023. Its best configuration ranks #485 of 372 on the BenchLeader Index at 42.0 ±8.0, in the lower half. It scores highest in coding (45) and lowest in maths (24). At $37.50 per million tokens blended it is among the most expensive ranked models. Last measured 2 Sept 2026.
- Blended price
- $37.50/M
- $30.00 in · $60.00 out
- Output speed
- –
- First answer
- –
- Context
- 8k
- Overall index42
- Reasoning35
- Coding45
- Maths24
- Human preference45
Versions
OpenAI has shipped 14 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.
| Model | Released | Index | Rank |
|---|---|---|---|
| GPT-5.6 | 9 Jul 2026 | – | – |
| GPT-5.5 | 23 Apr 2026 | 66.7 | #17 |
| GPT-5.5xhigh | 23 Apr 2026 | – | – |
| GPT-5.4 | 5 Mar 2026 | 63.5 | #46 |
| GPT-5.1 | 12 Nov 2025 | 59.6 | #100 |
| GPT-5 | 7 Aug 2025 | 59.9 | #97 |
| GPT-4.1 | 14 Apr 2025 | 47.4 | #358 |
| GPT-4.5 | 27 Feb 2025 | 49.8 | #310 |
| GPT 4 0125 | 25 Jan 2024 | 45.5 | #396 |
| GPT-4 | 6 Nov 2023 | 44.5 | #424 |
| GPT 4 0613 | 13 Jun 2023 | 39.2 | #549 |
| GPT 4 0314this page | 14 Mar 2023 | 42.0 | #485 |
| GPT-5.2 | – | 61.6 | #65 |
| GPT-3.5 | – | – | – |
Benchmark results
One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.
Reasoning
| Benchmark | default | Source |
|---|---|---|
| GPQA Diamond | 35.7%#248 | Epoch AI Benchmarking Hub |
| LMArena Hard Prompts | 1305#239 | LMArena |
Coding
| Benchmark | default | Source |
|---|---|---|
| LMArena Coding | 1329#243 | LMArena |
Maths
| Benchmark | default | Source |
|---|---|---|
| OTIS Mock AIME | 0.6%#258 | Epoch AI Benchmarking Hub |
Human preference
| Benchmark | default | Source |
|---|---|---|
| LMArena Text | 1287#249 | LMArena |
What a task costs
Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.
| Workload | Tokens in / out | Cost | With caching | Time |
|---|---|---|---|---|
| Chat reply | 400 / 300 | $0.030 | – | – |
| Summarise a 30-page report | 12,000 / 600 | $0.396 | – | – |
| Code edit | 6,000 / 1,500 | $0.270 | – | – |
| Agentic coding session | 60,000 / 4,000 | $2.04 | – | – |
| Structured extraction | 2,000 / 200 | $0.072 | – | – |
See also
Data as of 10 Sept 2026. Compare with another model.