BenchLeader

Agnes 3.0 Flash vs GPT-5.6 Sol

Verdict
  • GPT-5.6 Sol (max) leads on quality: 68.8 vs 62.3.
  • Agnes 3.0 Flash is stronger in agents & tools.
  • GPT-5.6 Sol (max) is stronger in composite, knowledge, long context, reasoning, coding, instruction following, maths, multimodal.
  • Agnes 3.0 Flash is 107× cheaper ($0.075 vs $8.00 per 1M blended).
MetricAgnes 3.0 FlashGPT-5.6 Sol (max)
BenchLeader Index62.368.8
Agents & tools score76.866.6
Composite score74.079.2
Knowledge score59.166.5
Long context score67.068.5
Reasoning score71.173.7
Coding score67.0
Instruction following score71.0
Maths score74.3
Multimodal score68.1
Blended price $/M$0.075$8.00
Output speed61 tok/s
Time to first answer127.3 s
Context window1M1.1M
GPQA Diamond93.5%
FrontierMath Tiers 1–389.1%
FrontierMath Tier 482.9%
OTIS Mock AIME100.0%
SimpleQA Verified69.7%
Terminal-Bench37.3%
OSWorld-Verified 2.027.3%
SciCode56.1%
WeirdML87.0%
APEX-Agents39.9%
ProofBench83.0%
LiveBench81.0%
LiveBench Reasoning91.7%
LiveBench Coding83.9%
LiveBench Agentic Coding56.2%
LiveBench Mathematics96.2%
LiveBench Data Analysis79.8%
LiveBench Language87.7%
LiveBench Instruction Following71.8%
AA Intelligence Index35.547.1
IFBench72.7%
AA-LCR81.0%84.0%
MMMU-Pro83.4%
AA-Omniscience-10.622.0
Terminal-Bench Hard65.9%
GPQA Diamond (AA)92.4%94.1%
Humanity's Last Exam (AA)38.5%49.5%
SciCode (AA)51.6%57.1%
τ²-Bench Telecom (AA)85.1%
LiveCodeBench82.6%
MMLU-Pro89.1%
IOI91.2%
LegalBench87.0%
CorpFin64.4%
TaxEval74.8%
Terminal-Bench 2.1 (Vals)85.8%
SWE-bench (Vals)96.2%
GPQA Diamond (Vals)95.2%
Vals Index63.7
PRBench Finance50.5%
PRBench Legal50.5%
ARC-AGI-196.5%
ARC-AGI-292.5%
ARC-AGI-37.8%
CritPt15.1%32.3%
GDPval (AA)53.7%54.3%
τ²-Bench Banking (AA)47.6%44.3%
ITBench SRE (AA)56.2%
Analyst Agent (AA)47.5%
BioMysteryBench71.1%
Code Migration52.9%
CUA-bench8.3%
Excel Modeling Benchmark72.3%
Finance Agent v253.8%
Harvey's Legal Agent Benchmark2.5%
Legal Research Bench48.1%
MedCode44.0%
MedScribe85.2%
MMMU-Pro (Vals)88.8%
MortgageTax67.3%
MysteryMechanism33.3%
ProgramBench1.5%
Public Benefits Bench66.5%
SAGE52.6%
SkillsBench54.1%
SREBench30.5%
Tax Agent Bench68.0%
Terminal-Bench 4.0 (Vals)27.8%
Terminal-Bench Science12.9%
Time Horizon Index: KSP23.8%
Vals Multimodal Index72.6%
Vibe Code Bench 1-10020.0%
Vibe Code Bench v1.180.5%
Web Search Index43.6%
DrugDiscoveryBench56.5%
Chess Puzzles55.0%
EBR-bench44.8%
Mystery Game Puzzles58.0%
PostTrainBench36.2%
DeepSWE72.7%
LMCA58.4%
DTBench95.5%
CursorBench67.2%
ALE-Bench2176.9
GDP.pdf30.7%
FrontierSWE32.2%
BTF-313.7%

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

Agnes 3.0 Flash vs GPT-5.6 Sol: questions

Is Agnes 3.0 Flash better than GPT-5.6 Sol?
GPT-5.6 Sol (max) leads on quality: 68.8 vs 62.3. The BenchLeader Index combines every independent quality benchmark; GPT-5.6 Sol (max) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is Agnes 3.0 Flash better than GPT-5.6 Sol for agentic tasks?
Agnes 3.0 Flash scores higher in agentic tasks (77 vs 67 on the category index, where 50 is average).
Which is cheaper, Agnes 3.0 Flash or GPT-5.6 Sol?
Agnes 3.0 Flash is cheaper: $0.075 against $8.00 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
GPT-5.6 Sol accepts more context: 1.1M against 1M tokens.