BenchLeader

Gemini 3.5 Flash vs GPT-5.6 Terra

Verdict
  • Gemini 3.5 Flash (medium) and GPT-5.6 Terra (xhigh) are level on quality (63.9 vs 64.2).
  • Gemini 3.5 Flash (medium) is stronger in human preference, instruction following, knowledge, multimodal, reasoning.
  • GPT-5.6 Terra (xhigh) is stronger in agents & tools, coding, composite, long context, maths.
  • Gemini 3.5 Flash (medium) is 1.3× cheaper ($3.38 vs $4.50 per 1M blended).
  • Gemini 3.5 Flash (medium) streams 2.2× faster (209 vs 93 tokens per second).
MetricGemini 3.5 Flash (medium)GPT-5.6 Terra (xhigh)
BenchLeader Index63.964.2
Agents & tools score67.869.6
Coding score58.963.7
Composite score71.477.1
Human preference score67.466.3
Instruction following score72.265.0
Knowledge score73.960.9
Long context score63.666.0
Multimodal score67.663.0
Reasoning score66.858.2
Maths score71.0
Blended price $/M$3.38$4.50
Output speed209 tok/s93 tok/s
Time to first answer17.0 s32.5 s
Context window1M1M
SimpleBench48.9%
SciCode51.6%
ProofBench74.0%
LMArena Text14761466
LMArena Hard Prompts14931491
LMArena Coding15031519
LMArena WebDev14911521
LMArena Vision13061269
LMArena Agent1.5
AA Intelligence Index33.638.2
IFBench74.6%66.3%
AA-LCR74.3%79.0%
MMMU-Pro83.9%79.5%
AA-Omniscience20.8-3.0
Terminal-Bench Hard39.4%62.9%
GPQA Diamond (AA)92.1%90.8%
Humanity's Last Exam (AA)41.3%41.9%
SciCode (AA)52.3%
τ²-Bench Telecom (AA)95.6%80.4%
LiveCodeBench85.9%
MMLU-Pro86.7%
IOI65.3%
CorpFin65.3%
TaxEval76.2%
GPQA Diamond (Vals)90.9%
ARC-AGI-194.0%
ARC-AGI-274.2%
ARC-AGI-30.7%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.