BenchLeader

Gemini 3.5 Flash vs Gemini 3 Flash

Verdict
  • Gemini 3.5 Flash (medium) leads on quality: 64.1 vs 61.6.
  • Gemini 3.5 Flash (medium) is stronger in agents & tools, coding, composite, human preference, knowledge, multimodal, reasoning.
  • Gemini 3 Flash (thinking) is stronger in instruction following, long context.
  • Gemini 3 Flash (thinking) is 3.0× cheaper ($1.13 vs $3.38 per 1M blended).
MetricGemini 3.5 Flash (medium)Gemini 3 Flash (thinking)
BenchLeader Index64.161.6
Agents & tools score68.067.4
Coding score58.9
Composite score71.562.3
Human preference score67.8
Instruction following score72.575.4
Knowledge score74.069.0
Long context score63.665.5
Multimodal score67.664.4
Reasoning score67.1
Blended price $/M$3.38$1.13
Output speed217 tok/s180 tok/s
Time to first answer15.5 s6.1 s
Context window1M1M
LMArena Text1476
LMArena Hard Prompts1493
LMArena Coding1503
LMArena WebDev1491
LMArena Vision1306
AA Intelligence Index33.626.3
IFBench74.6%78.0%
AA-LCR74.3%78.0%
MMMU-Pro83.9%79.9%
AA-Omniscience20.810.1
Terminal-Bench Hard39.4%38.6%
GPQA Diamond (AA)92.1%89.8%
Humanity's Last Exam (AA)41.3%36.6%
τ²-Bench Telecom (AA)95.6%80.4%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Gemini 3.5 Flash vs Gemini 3 Flash: questions

Is Gemini 3.5 Flash better than Gemini 3 Flash?
Gemini 3.5 Flash (medium) leads on quality: 64.1 vs 61.6. The BenchLeader Index combines every independent quality benchmark; Gemini 3.5 Flash (medium) is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Gemini 3.5 Flash better than Gemini 3 Flash for agentic tasks?
Gemini 3.5 Flash scores higher in agentic tasks (68 vs 67 on the category index, where 50 is average).
Which is cheaper, Gemini 3.5 Flash or Gemini 3 Flash?
Gemini 3 Flash is cheaper: $1.13 against $3.38 per million tokens, blended at three input tokens per output token.
Which is faster, Gemini 3.5 Flash or Gemini 3 Flash?
Gemini 3.5 Flash streams faster: 217 against 180 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.