BenchLeader

GPT-6.1 Sol vs GPT-6 Sol

Verdict
  • GPT-6.1 Sol (max) leads on quality: 69.3 vs 65.6.
  • GPT-6.1 Sol (max) is stronger in agents & tools, coding, composite, human preference, knowledge, maths, multimodal, reasoning.
  • GPT-6 Sol (max) is stronger in long context.
  • They cost about the same ($4.00 per 1M blended).
  • GPT-6 Sol (max) streams 1.6× faster (86 vs 56 tokens per second).
MetricGPT-6.1 Sol (max)GPT-6 Sol (max)
BenchLeader Index69.365.6
Agents & tools score63.960.3
Coding score72.467.6
Composite score79.073.1
Human preference score68.164.7
Knowledge score78.870.1
Long context score66.867.1
Maths score72.569.5
Multimodal score66.261.9
Reasoning score76.471.8
Blended price $/M$4.00$4.00
Output speed56 tok/s86 tok/s
Time to first answer326.9 s130.4 s
Context window1.1M1.1M
GPQA Diamond95.4%94.3%
FrontierMath Tiers 1–393.7%89.8%
FrontierMath Tier 4100.0%90.0%
OTIS Mock AIME100.0%100.0%
SimpleQA Verified73.9%60.7%
Terminal-Bench58.2%49.4%
SciCode54.2%57.6%
APEX-Agents60.0%54.3%
FrontierCode–49.3%
LMArena Text14841456
LMArena Hard Prompts15071482
LMArena Coding15451519
LMArena WebDev17551688
LMArena Vision12881245
LMArena Agent11.79.9
LiveBench81.6%79.3%
LiveBench Reasoning92.6%88.7%
LiveBench Coding80.4%81.8%
LiveBench Agentic Coding54.5%52.9%
LiveBench Mathematics96.8%96.4%
LiveBench Data Analysis82.7%81.2%
LiveBench Language90.1%85.3%
LiveBench Instruction Following74.2%68.6%
AA Intelligence Index v4.3.251.847.6
AA-LCR83.0%83.7%
MMMU-Pro86.0%83.0%
AA-Omniscience41.527.1
Humanity's Last Exam (AA)52.9%47.9%
SciCode (AA)54.2%57.6%
IOI96.9%82.6%
Terminal-Bench 2.1 (Vals)–83.2%
Vals Index61.157.5
ARC-AGI-196.5%95.5%
ARC-AGI-294.2%89.6%
ARC-AGI-396.2%23.0%
CritPt31.7%30.9%
GDPval-AA v2.153.8%50.4%
ITBench SRE (AA)–49.4%
Analyst Agent (AA)50.0%–
BioMysteryBench79.6%74.8%
Code Migration65.1%57.2%
CyberBench39.3%78.0%
Excel Modeling Benchmark70.8%71.5%
Finance Agent v252.0%49.0%
Harvey's Legal Agent Benchmark5.4%1.7%
Legal Research Bench38.5%28.9%
MedCode48.8%47.1%
MedScribe86.5%82.0%
MysteryMechanism46.4%30.2%
ProgramBench–2.0%
Public Benefits Bench59.3%56.6%
SAGE46.5%44.8%
SREBench50.8%–
Tax Agent Bench62.3%53.0%
Terminal-Bench 4.0 (Vals)55.0%44.4%
Terminal-Bench Science–30.0%
Vibe Code Bench v1.188.9%87.8%
LMArena Maths14881446
LMArena Creative Writing14621436
LMArena Instruction Following14881451
LMArena Multi-turn14871467
LMArena Longer Queries14961463
Chess Puzzles61.0%–
EBR-bench54.3%53.3%
Mystery Game Puzzles80.0%56.0%
LMCA–59.1%
DTBench–97.3%
ALE-Bench–2462
GDP.pdf–26.4%
Terminal-Bench 4.0 (AA)56.1%43.9%
AutomationBench64.9%61.6%
GDP.pdf31.0%25.2%
MLCR33.9%16.1%
Harvey LAB6.9%3.6%
AA-Omniscience: accuracy62.1%54.5%
AA-Omniscience: non-hallucination45.7%39.9%
AA-Briefcase v1.115571480

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

GPT-6.1 Sol vs GPT-6 Sol: questions

Is GPT-6.1 Sol better than GPT-6 Sol?
GPT-6.1 Sol (max) leads on quality: 69.3 vs 65.6. The BenchLeader Index combines every independent quality benchmark; GPT-6.1 Sol (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is GPT-6.1 Sol better than GPT-6 Sol for coding?
GPT-6.1 Sol scores higher in coding (72 vs 68 on the category index, where 50 is average).
Is GPT-6.1 Sol better than GPT-6 Sol for agentic tasks?
GPT-6.1 Sol scores higher in agentic tasks (64 vs 60 on the category index, where 50 is average).
Which is cheaper, GPT-6.1 Sol or GPT-6 Sol?
GPT-6.1 Sol is cheaper: $4.00 against $4.00 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-6.1 Sol or GPT-6 Sol?
GPT-6 Sol streams faster: 86 against 56 output tokens per second.
Which has the larger context window?
Both accept 1.1M tokens of context.