BenchLeader

GPT-6 Sol vs o3

Verdict
  • GPT-6 Sol (max) leads on quality: 67.7 vs 60.0.
  • GPT-6 Sol (max) is stronger in composite, knowledge, long context, multimodal, reasoning.
  • o3 is stronger in agents & tools, coding, human preference, instruction following, maths.
  • They cost about the same ($3.50 per 1M blended).
MetricGPT-6 Sol (max)o3
BenchLeader Index67.760.0
Composite score86.253.4
Knowledge score75.456.7
Long context score67.763.0
Multimodal score67.157.5
Reasoning score95.052.5
Agents & tools score70.2
Coding score61.9
Human preference score62.3
Instruction following score65.1
Maths score71.2
Blended price $/M$4.00$3.50
Output speed126 tok/s118 tok/s
Time to first answer107.2 s8.2 s
Context window1.1M200k
Epoch Capabilities Index146.9
LMArena Text1432
LMArena Hard Prompts1441
LMArena Coding1460
LMArena Vision1214
AA Intelligence Index47.520.2
IFBench71.4%
AA-LCR83.7%74.7%
MMMU-Pro83.3%70.1%
AA-Omniscience27.1-15.6
Terminal-Bench Hard37.1%
GPQA Diamond (AA)82.7%
Humanity's Last Exam (AA)47.9%20.1%
SciCode (AA)57.6%
τ²-Bench Telecom (AA)80.7%
PRBench Finance47.7%
PRBench Legal48.6%
MMMU (validation)82.9%
MMMU-Pro (official)76.4%
Kagi LLM Benchmark67.6%
IFEval (HELM)86.9%
Omni-MATH (HELM)71.4%
WildBench (HELM)86.1%
MMLU-Pro (HELM)85.9%
GPQA Diamond (HELM)75.3%
HELM Capabilities mean81.1%
Aider Polyglot81.3%
SWE-bench Verified (bash only)58.4%
SWE-bench Verified (any scaffold)58.4%
BFCL Overall63.0%
CritPt30.9%1.1%
GDPval (AA)49.4%
FORTRESS16.0%
PropensityBench10.5%
SciPredict17.9%
VTB13.7%
LMArena Maths1448
LMArena Creative Writing1382
LMArena Instruction Following1402
LMArena Multi-turn1420
LMArena Longer Queries1410
METR Time Horizons63.6%
ForecastBench62.5%
Terminal-Bench 4.0 (AA)43.9%
AutomationBench61.6%
GDP.pdf24.8%
MLCR16.1%
AA-Omniscience: accuracy54.5%38.6%
AA-Omniscience: non-hallucination39.9%11.9%
AA-Briefcase1483
BrowseComp-Plus50.5%

Data as of 2026-09-23. Best configuration of each model; every score links to its source on the model pages.

GPT-6 Sol vs o3: questions

Is GPT-6 Sol better than o3?
GPT-6 Sol (max) leads on quality: 67.7 vs 60.0. The BenchLeader Index combines every independent quality benchmark; GPT-6 Sol (max) is ahead overall as of 2026-09-23, but check the category scores for your use.
Which is cheaper, GPT-6 Sol or o3?
o3 is cheaper: $3.50 against $4.00 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-6 Sol or o3?
GPT-6 Sol streams faster: 126 against 118 output tokens per second.
Which has the larger context window?
GPT-6 Sol accepts more context: 1.1M against 200k tokens.