BenchLeader

GPT-6.1 Sol

Reasoning effort

GPT-6.1 Sol is an OpenAI proprietary reasoning model, released 29 Sept 2026. Its best configuration (max reasoning effort) ranks #10 of 760 on the BenchLeader Index at 69.3 ±3.2, in the top ten. The ± is the point: 48 other configurations score within that range, so they and this one cannot be told apart on quality alone — price and speed are what separate them. It scores highest in composite (79) and lowest in agents & tools (64). At $4.00 per million tokens blended it is among the most expensive ranked models. Output speed of 56 tokens per second puts it slower than most, with a first answer in 326.9 s. It has been measured at 6 reasoning-effort settings; this summary describes the best-scoring one, and the tabs above switch between them. Last measured 9 Oct 2026.

Blended price
$4.00/M
$2.00 in · $10.00 out
Output speed
56 tok/s
measured by Artificial Analysis
First answer
327 s
first token 5.01 s
Context
1.1M
Full answer
336 s
median, reasoning included
Cost per run
$0.724
one full Intelligence Index run
Released
29 Sept 2026
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index69
  2. Reasoning76
  3. Coding72
  4. Agents & tools64
  5. Maths73
  6. Knowledge79
  7. Human preference68
  8. Multimodal66
  9. Long context67
  10. Composite79

Versions

OpenAI has shipped 3 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
GPT-6.1 Solmaxthis page29 Sept 202669.3#10
GPT-6 Solmax22 Sept 202665.6#35
GPT-5.6 Solmax9 Jul 202667.5#22

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Not stated” collects results from publishers that did not say which setting they used; for a reasoning model that is usually its thinking mode, but we do not assume it. Pick a setting here or at the top of the page to see its price, speed and category scores.

EffortIndexRankSpeedFirst answerChat reply costCategories
low65.2#4047 tok/s2.80 s$0.0038Coding 59 · Composite 77 · Knowledge 79 · Long context 67 · Multimodal 66 · Reasoning 73
medium67.5#2449 tok/s5.76 s$0.0038Coding 62 · Composite 84 · Knowledge 80 · Long context 67 · Multimodal 67 · Reasoning 77
high67.9#1650 tok/s60 s$0.0038Coding 63 · Composite 87 · Knowledge 80 · Long context 66 · Multimodal 68 · Reasoning 79
xhigh67.2#2553 tok/s152 s$0.0038Coding 63 · Composite 78 · Knowledge 80 · Long context 65 · Multimodal 68 · Reasoning 80
maxbest69.3#1056 tok/s327 s$0.0038Agents & tools 64 · Coding 72 · Composite 79 · Human preference 68 · Knowledge 79 · Long context 67 · Maths 73 · Multimodal 66 · Reasoning 76
not stated––35 tok/s5.01 s$0.0038Instruction following 83 · Knowledge 56 · Maths 79

What more thinking costs

Turning the effort up buys index points and multiplies the bill. Here max costs 5.5× low for +4.1 on the index — 1.7 points per doubling of spend. Compare that with other models.

low $0.131max $0.724

How its index has moved

2 Oct 2026 to 9 Oct 2026
646873
  • max
  • high
  • medium
  • xhigh

The index is recomputed from scratch every day, so a line moves when a new benchmark result lands, when a publisher revises a score, or when the models it is normalised against change. Early movement usually means the score is still settling.

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
GPQA Diamond––––95.4%#3–Epoch AI Benchmarking Hub
LMArena Hard Prompts––––1507#19–LMArena
LiveBench Reasoningnot in index–––91.6%#692.6%#2–LiveBench
Humanity's Last Exam (AA)not in index47.4%#3649.9%#2451.4%#2152.6%#2052.9%#17–Artificial Analysis
ARC-AGI-193.5%#4495.5%#2998.5%#198.5%#196.5%#21–ARC Prize
ARC-AGI-276.7%#4486.7%#2491.7%#991.7%#994.2%#2–ARC Prize
ARC-AGI-382.8%#1191.0%#1095.0%#996.4%#796.2%#8–ARC Prize
CritPt24.9%#4227.7%#2730.0%#1531.7%#231.7%#2–Artificial Analysis
MysteryMechanismnot in index––––46.4%#5–Vals AI
Chess Puzzlesnot in index––––61.0%#4–Epoch AI Benchmarking Hub
Mystery Game Puzzlesnot in index––––80.0%#2–Epoch AI Benchmarking Hub

Coding

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
SciCode53.2%#7353.2%#7355.8%#4055.7%#4154.2%#59–SciCode
FrontierCode–50.2%#7––––Cognition
LMArena Coding––––1545#8–LMArena
LMArena WebDev––––1755#4–LMArena
LiveBench Codingnot in index–––80.7%#1780.4%#19–LiveBench
SciCode (AA)not in index53.2%#7353.2%#7355.8%#4055.7%#4254.2%#61–Artificial Analysis
IOI––––96.9%#3–Vals AI
Code Migrationnot in index––––65.1%#5–Vals AI
Vibe Code Bench v1.1not in index––––88.9%#7–Vals AI
Terminal-Bench 4.0 (AA)not in index30.8%#4148.0%#1951.5%#1654.0%#1156.1%#9–Artificial Analysis

Agents & tools

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
Terminal-Bench––––58.2%#17–Terminal-Bench
APEX-Agents––––60.0%#12–Mercor
LMArena Agent––––11.7#5–LMArena
LiveBench Agentic Codingnot in index–––56.8%#2454.5%#31–LiveBench
GDPval-AA v2.1not in index39.9%#9746.6%#6249.3%#4850.5%#4453.8%#35–Artificial Analysis
Analyst Agent (AA)not in index––––50.0%#7–Artificial Analysis
CyberBenchnot in index––––39.3%#39–Vals AI
Finance Agent v2not in index––––52.0%#30–Vals AI
Harvey's Legal Agent Benchmarknot in index––––5.4%#28–Vals AI
Legal Research Benchnot in index––––38.5%#27–Vals AI
SREBenchnot in index––––50.8%#2–Vals AI
Tax Agent Benchnot in index––––62.3%#26–Vals AI
Terminal-Bench 4.0 (Vals)not in index––––55.0%#6–Vals AI
AutomationBenchnot in index––64.5%#1666.6%#1064.9%#14–Artificial Analysis
GDP.pdfnot in index––32.0%#231.8%#331.0%#4–Artificial Analysis
Harvey LABnot in index––––6.9%#4–Harvey
AA-Briefcase v1.1not in index––1465#291503#251557#16–Artificial Analysis

Maths

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
FrontierMath Tiers 1–3––––93.7%#1–Epoch AI Benchmarking Hub
FrontierMath Tier 4––––100.0%#1–Epoch AI Benchmarking Hub
OTIS Mock AIME––––100.0%#1–Epoch AI Benchmarking Hub
ProofBencheffort not stated–––––99.0%#4Vals AI
LiveBench Mathematicsnot in index–––96.5%#796.8%#3–LiveBench
LMArena Maths––––1488#29–LMArena

Knowledge

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
SimpleQA Verified––––73.9%#2–Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index–––82.2%#382.7%#2–LiveBench
AA-Omniscience37.6#2040.0#1841.5#1240.9#1341.5#11–Artificial Analysis
PRBench Financeeffort not stated–––––46.2%#17Scale AI SEAL
PRBench Legaleffort not stated–––––47.9%#15Scale AI SEAL
BioMysteryBenchnot in index––––79.6%#2–Vals AI
Excel Modeling Benchmarknot in index––––70.8%#12–Vals AI
MedCodenot in index––––48.8%#26–Vals AI
MedScribenot in index––––86.5%#21–Vals AI
Public Benefits Benchnot in index––––59.3%#30–Vals AI
EBR-benchnot in index––––54.3%#4–Epoch AI Benchmarking Hub
AA-Omniscience: accuracynot in index58.9%#2560.4%#1960.8%#1760.8%#1662.1%#12–Artificial Analysis
AA-Omniscience: non-hallucinationnot in index48.4%#12048.4%#12150.6%#10749.1%#11445.7%#135–Artificial Analysis

Instruction following

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
LiveBench Languagenot in index–––88.7%#690.1%#2–LiveBench
LiveBench Instruction Followingnot in index–––71.3%#2474.2%#12–LiveBench
MultiChallengeeffort not stated–––––82.0%#1Scale AI SEAL
LMArena Instruction Followingnot in index––––1488#15–LMArena

Human preference

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
LMArena Text––––1484#21–LMArena
LMArena Creative Writingnot in index––––1462#26–LMArena
LMArena Multi-turnnot in index––––1487#25–LMArena
LMArena Longer Queriesnot in index––––1496#20–LMArena

Multimodal

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
LMArena Vision––––1288#30–LMArena
MMMU-Pro83.1%#2783.9%#2384.9%#1385.1%#1186.0%#6–Artificial Analysis
SAGEnot in index––––46.5%#35–Vals AI

Long context

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
AA-LCR84.0%#1183.3%#1982.3%#3779.7%#8383.0%#24–Artificial Analysis
MLCRnot in index––––33.9%#15–Artificial Analysis

Composite

Benchmarklowmediumhighxhighmaxnot statedSourceTrend
Epoch Capabilities Indexnot in indexeffort not stated–––––166.1#3Epoch AI Benchmarking Hub
LiveBench–––81.1%#881.6%#6–LiveBench
AA Intelligence Index v4.3.242.1#4947.8#2450.2#1751.0#1451.8#11–Artificial Analysis
Vals Indexnot in index––––61.1#8–Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
OpenAI84 tok/s2.35 s$12.00$60.001.1M–
Amazon Bedrock46 tok/s1.92 s$2.20$11.001.1M–
Azure45 tok/s1.93 s$2.20$11.001.1M–
OpenAI44 tok/s2.56 s$4.00$20.001.1M–
Azure43 tok/s4.92 s$2.00$10.001.1M–
OpenAI40 tok/s2.77 s$2.00$10.001.1M–
Azure37 tok/s3.71 s$2.20$11.001.1M–
OpenAI27 tok/s9.96 s$1.00$5.001.1M–

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.100 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0038$0.00325.5 min
Summarise a 30-page report12,000 / 600$0.030$0.0135.6 min
Code edit6,000 / 1,500$0.027$0.0185.9 min
Agentic coding session60,000 / 4,000$0.160$0.0746.7 min
Structured extraction2,000 / 200$0.0060$0.00325.5 min

See also

Data as of 11 Oct 2026. Compare these configurations.

Cite as: BenchLeader, “GPT-6.1 Sol: benchmarks, pricing, speed and rank”, https://www.benchleader.com/models/gpt-6-1-sol, data as of 11 Oct 2026.