BenchLeader

Gemini 3.7 Flash vs GPT-5.6 Sol

Verdict
  • GPT-5.6 Sol (high) leads on quality: 68.0 vs 64.4.
  • Gemini 3.7 Flash (medium) is stronger in knowledge, long context, multimodal, reasoning.
  • GPT-5.6 Sol (high) is stronger in coding, composite, agents & tools, instruction following.
  • Gemini 3.7 Flash (medium) is 5.3× cheaper ($1.50 vs $8.00 per 1M blended).
  • Gemini 3.7 Flash (medium) streams 4.8× faster (282 vs 59 tokens per second).
MetricGemini 3.7 Flash (medium)GPT-5.6 Sol (high)
BenchLeader Index64.468.0
Coding score68.372.0
Composite score78.982.5
Knowledge score75.273.7
Long context score68.167.4
Multimodal score69.466.5
Reasoning score63.562.9
Agents & tools score87.6
Instruction following score67.6
Blended price $/M$1.50$8.00
Output speed282 tok/s59 tok/s
Time to first answer5.3 s31.4 s
Context window1M1M
SciCode57.9%56.9%
WeirdML88.8%
AA Intelligence Index39.642.5
IFBench69.2%
AA-LCR83.0%81.7%
MMMU-Pro84.7%81.8%
AA-Omniscience23.720.4
Terminal-Bench Hard62.1%
GPQA Diamond (AA)92.1%92.8%
Humanity's Last Exam (AA)39.0%46.0%
SciCode (AA)59.8%57.8%
τ²-Bench Telecom (AA)83.3%
EnigmaEval37.1%
ARC-AGI-191.2%97.0%
ARC-AGI-263.8%85.4%
ARC-AGI-32.2%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.