BenchLeader

Gemini 3.5 Flash vs GPT-5.6 Sol

Verdict
  • GPT-5.6 Sol (high) leads on quality: 68.0 vs 63.9.
  • Gemini 3.5 Flash (medium) is stronger in human preference, instruction following, knowledge, multimodal, reasoning.
  • GPT-5.6 Sol (high) is stronger in agents & tools, coding, composite, long context.
  • Gemini 3.5 Flash (medium) is 2.4× cheaper ($3.38 vs $8.00 per 1M blended).
  • Gemini 3.5 Flash (medium) streams 3.5× faster (209 vs 59 tokens per second).
MetricGemini 3.5 Flash (medium)GPT-5.6 Sol (high)
BenchLeader Index63.968.0
Agents & tools score67.887.6
Coding score58.972.0
Composite score71.482.5
Human preference score67.4
Instruction following score72.267.6
Knowledge score73.973.7
Long context score63.667.4
Multimodal score67.666.5
Reasoning score66.862.9
Blended price $/M$3.38$8.00
Output speed209 tok/s59 tok/s
Time to first answer17.0 s31.4 s
Context window1M1M
SciCode56.9%
WeirdML88.8%
LMArena Text1476
LMArena Hard Prompts1493
LMArena Coding1503
LMArena WebDev1491
LMArena Vision1306
AA Intelligence Index33.642.5
IFBench74.6%69.2%
AA-LCR74.3%81.7%
MMMU-Pro83.9%81.8%
AA-Omniscience20.820.4
Terminal-Bench Hard39.4%62.1%
GPQA Diamond (AA)92.1%92.8%
Humanity's Last Exam (AA)41.3%46.0%
SciCode (AA)57.8%
τ²-Bench Telecom (AA)95.6%83.3%
EnigmaEval37.1%
ARC-AGI-197.0%
ARC-AGI-285.4%
ARC-AGI-32.2%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.