BenchLeader

Gemini 3.5 Flash vs GPT-5.2-Codex

Verdict
  • Gemini 3.5 Flash (medium) leads on quality: 64.1 vs 60.8.
  • Gemini 3.5 Flash (medium) is stronger in agents & tools, coding, composite, human preference, knowledge, multimodal, reasoning.
  • GPT-5.2-Codex is stronger in instruction following, long context.
  • Gemini 3.5 Flash (medium) is 1.4× cheaper ($3.38 vs $4.81 per 1M blended).
  • Gemini 3.5 Flash (medium) streams 8.7× faster (217 vs 25 tokens per second).
MetricGemini 3.5 Flash (medium)GPT-5.2-Codex
BenchLeader Index64.160.8
Agents & tools score68.063.5
Coding score58.955.7
Composite score71.556.8
Human preference score67.8
Instruction following score72.575.1
Knowledge score74.063.1
Long context score63.667.8
Multimodal score67.660.7
Reasoning score67.1
Blended price $/M$3.38$4.81
Output speed217 tok/s25 tok/s
Time to first answer15.5 s6.7 s
Context window1M400k
Terminal-Bench66.5%
APEX-Agents27.6%
LMArena Text1476
LMArena Hard Prompts1493
LMArena Coding1503
LMArena WebDev14911339
LMArena Vision1306
LiveBench74.0%
LiveBench Reasoning77.7%
LiveBench Coding83.6%
LiveBench Agentic Coding49.4%
LiveBench Mathematics88.8%
LiveBench Data Analysis78.2%
LiveBench Language73.7%
AA Intelligence Index33.628.5
IFBench74.6%77.6%
AA-LCR74.3%82.3%
MMMU-Pro83.9%76.3%
AA-Omniscience20.8-2.2
Terminal-Bench Hard39.4%37.1%
GPQA Diamond (AA)92.1%89.9%
Humanity's Last Exam (AA)41.3%35.7%
τ²-Bench Telecom (AA)95.6%92.1%
LiveCodeBench88.0%
SWE-Bench Pro41.0%
SWE-bench Verified (bash only)72.8%
SWE-bench Verified (any scaffold)72.8%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Gemini 3.5 Flash vs GPT-5.2-Codex: questions

Is Gemini 3.5 Flash better than GPT-5.2-Codex?
Gemini 3.5 Flash (medium) leads on quality: 64.1 vs 60.8. The BenchLeader Index combines every independent quality benchmark; Gemini 3.5 Flash (medium) is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Gemini 3.5 Flash better than GPT-5.2-Codex for coding?
Gemini 3.5 Flash scores higher in coding (59 vs 56 on the category index, where 50 is average).
Is Gemini 3.5 Flash better than GPT-5.2-Codex for agentic tasks?
Gemini 3.5 Flash scores higher in agentic tasks (68 vs 64 on the category index, where 50 is average).
Which is cheaper, Gemini 3.5 Flash or GPT-5.2-Codex?
Gemini 3.5 Flash is cheaper: $3.38 against $4.81 per million tokens, blended at three input tokens per output token.
Which is faster, Gemini 3.5 Flash or GPT-5.2-Codex?
Gemini 3.5 Flash streams faster: 217 against 25 output tokens per second.
Which has the larger context window?
Gemini 3.5 Flash accepts more context: 1M against 400k tokens.