BenchLeader
OpenAIReasoning modelNewReleased 22 Sept 2026

GPT-6 Sol

Reasoning effort

GPT-6 Sol is an OpenAI proprietary reasoning model, released 22 Sept 2026. Its best configuration (max reasoning effort) ranks #22 of 432 on the BenchLeader Index at 67.7 ±9.5, in the top 25. It scores highest in reasoning (95) and lowest in multimodal (67). At $4.00 per million tokens blended it is among the most expensive ranked models. Output speed of 126 tokens per second puts it faster than most, with a first answer in 107.2 s. It has been measured at 6 reasoning-effort settings; this summary describes the best-scoring one, and the tabs above switch between them. Last measured 22 Sept 2026.

Blended price
$4.00/M
$2.00 in · $10.00 out
Output speed
126 tok/s
measured by Artificial Analysis
First answer
107 s
first token 2.77 s
Context
1.1M
Full answer
111 s
median, reasoning included
Cost per run
$1.06
one full Intelligence Index run
Released
22 Sept 2026
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index68
  2. Reasoning95
  3. Knowledge75
  4. Multimodal67
  5. Long context68
  6. Composite86

Versions

OpenAI has shipped 2 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
NewGPT-6 Solmaxthis page22 Sept 202667.7#22
GPT-5.6 Solmax9 Jul 202668.4#19

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Not stated” collects results from publishers that did not say which setting they used; for a reasoning model that is usually its thinking mode, but we do not assume it. Pick a setting here or at the top of the page to see its price, speed and category scores.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning54.1#249116 tok/s0.97 s$0.0038Composite 63 · Knowledge 63 · Long context 58 · Multimodal 52 · Reasoning 49
low61.7#82124 tok/s1.44 s$0.0038Composite 70 · Knowledge 75 · Long context 65 · Multimodal 63 · Reasoning 71
medium64.9#43114 tok/s1.96 s$0.0038Composite 77 · Knowledge 75 · Long context 67 · Multimodal 64 · Reasoning 85
high65.7#36123 tok/s7.57 s$0.0038Composite 81 · Knowledge 75 · Long context 68 · Multimodal 65 · Reasoning 87
xhigh66.4#33134 tok/s45 s$0.0038Composite 82 · Knowledge 75 · Long context 67 · Multimodal 66 · Reasoning 91
maxbest67.7#22126 tok/s107 s$0.0038Composite 86 · Knowledge 75 · Long context 68 · Multimodal 67 · Reasoning 95

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowmediumhighxhighmaxSource
Humanity's Last Exam (AA)not in index18.4%#20034.9%#9841.0%#6344.1%#3946.3%#3347.9%#25Artificial Analysis
CritPt4.0%#13816.3%#6824.6%#3625.4%#3328.0%#2130.9%#7Artificial Analysis

Coding

Benchmarkno reasoninglowmediumhighxhighmaxSource
SciCode (AA)not in index47.3%#10250.2%#8653.8%#5954.9%#4855.1%#4357.6%#21Artificial Analysis
Terminal-Bench 4.0 (AA)not in index13.1%#519.1%#6418.7%#4326.3%#3230.3%#3043.9%#16Artificial Analysis

Agents & tools

Benchmarkno reasoninglowmediumhighxhighmaxSource
GDPval (AA)not in index36.3%#8733.8%#9741.0%#7043.8%#5646.8%#4649.4%#35Artificial Analysis
AutomationBenchnot in index61.6%#14Artificial Analysis
GDP.pdfnot in index24.8%#15Artificial Analysis
AA-Briefcasenot in index1483#26Artificial Analysis

Knowledge

Benchmarkno reasoninglowmediumhighxhighmaxSource
AA-Omniscience-0.8#11626.5#3827.0#3526.8#3626.7#3727.1#34Artificial Analysis
AA-Omniscience: accuracynot in index45.2%#7651.3%#4953.5%#4053.7%#3853.9%#3754.5%#35Artificial Analysis
AA-Omniscience: non-hallucinationnot in index16.0%#29949.3%#10243.2%#12441.9%#12741.1%#13039.9%#135Artificial Analysis

Multimodal

Benchmarkno reasoninglowmediumhighxhighmaxSource
MMMU-Pro68.2%#15278.8%#5680.6%#3881.2%#3382.4%#2683.3%#22Artificial Analysis

Long context

Benchmarkno reasoninglowmediumhighxhighmaxSource
AA-LCR64.0%#22879.3%#8182.3%#3183.7%#1481.3%#4583.7%#14Artificial Analysis
MLCRnot in index16.1%#24Artificial Analysis

Composite

Benchmarkno reasoninglowmediumhighxhighmaxSource
AA Intelligence Index28.1#11133.9#7439.8#4642.8#3644.1#3147.5#18Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Amazon Bedrock88 tok/s1.96 s$2.20$11.001.1M
Azure73 tok/s4.94 s$2.20$11.001.1M
OpenAI67 tok/s3.03 s$1.00$5.001.1M
OpenAI53 tok/s2.77 s$2.00$10.001.1M
OpenAI38 tok/s3.22 s$4.00$20.001.1M
Azure37 tok/s4.99 s$2.00$10.001.1M
Azure12 tok/s1.73 s$2.20$11.001.1M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.200 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0038$0.00331.8 min
Summarise a 30-page report12,000 / 600$0.030$0.0141.9 min
Code edit6,000 / 1,500$0.027$0.0192.0 min
Agentic coding session60,000 / 4,000$0.160$0.0792.3 min
Structured extraction2,000 / 200$0.0060$0.00331.8 min

See also

Data as of 23 Sept 2026. Compare these configurations.

Cite as: BenchLeader, “GPT-6 Sol: benchmarks, pricing, speed and rank”, https://www.benchleader.com/models/gpt-6-sol, data as of 23 Sept 2026.