BenchLeader

Agnes 2.5 Pro Beta vs GPT-6 Sol

Verdict
  • GPT-6 Sol (max) leads on quality: 67.7 vs 60.0.
  • Agnes 2.5 Pro Beta is stronger in agents & tools.
  • GPT-6 Sol (max) is stronger in composite, knowledge, long context, multimodal, reasoning.
  • Agnes 2.5 Pro Beta is 27× cheaper ($0.150 vs $4.00 per 1M blended).
MetricAgnes 2.5 Pro BetaGPT-6 Sol (max)
BenchLeader Index60.067.7
Agents & tools score64.4
Composite score71.586.2
Knowledge score58.075.4
Long context score67.367.7
Multimodal score59.067.1
Reasoning score69.595.0
Blended price $/M$0.150$4.00
Output speed126 tok/s
Time to first answer107.2 s
Context window1M1.1M
AA Intelligence Index35.247.5
AA-LCR83.0%83.7%
MMMU-Pro75.4%83.3%
AA-Omniscience-10.527.1
GPQA Diamond (AA)90.5%
Humanity's Last Exam (AA)37.5%47.9%
SciCode (AA)48.8%57.6%
CritPt15.7%30.9%
GDPval (AA)40.6%49.4%
τ²-Bench Banking (AA)35.7%
Terminal-Bench 4.0 (AA)43.9%
Terminal-Bench 2.1 (AA)69.7%
AutomationBench61.6%
GDP.pdf24.8%
MLCR16.1%
AA-Omniscience: accuracy16.8%54.5%
AA-Omniscience: non-hallucination67.2%39.9%
AA-Briefcase1483

Data as of 2026-09-23. Best configuration of each model; every score links to its source on the model pages.

Agnes 2.5 Pro Beta vs GPT-6 Sol: questions

Is Agnes 2.5 Pro Beta better than GPT-6 Sol?
GPT-6 Sol (max) leads on quality: 67.7 vs 60.0. The BenchLeader Index combines every independent quality benchmark; GPT-6 Sol (max) is ahead overall as of 2026-09-23, but check the category scores for your use.
Which is cheaper, Agnes 2.5 Pro Beta or GPT-6 Sol?
Agnes 2.5 Pro Beta is cheaper: $0.150 against $4.00 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
GPT-6 Sol accepts more context: 1.1M against 1M tokens.