BenchLeader

Gemini 3.5 Flash vs Grok 4.5

Verdict
  • Gemini 3.5 Flash (medium) and Grok 4.5 are level on quality (63.9 vs 63.4).
  • Gemini 3.5 Flash (medium) is stronger in agents & tools, composite, human preference, instruction following, multimodal.
  • Grok 4.5 is stronger in coding, knowledge, long context, reasoning.
  • They cost about the same ($3.00 per 1M blended).
  • Gemini 3.5 Flash (medium) streams 3.8× faster (209 vs 56 tokens per second).
MetricGemini 3.5 Flash (medium)Grok 4.5
BenchLeader Index63.963.4
Agents & tools score67.858.6
Coding score58.959.9
Composite score71.466.0
Human preference score67.466.9
Instruction following score72.2
Knowledge score73.976.0
Long context score63.666.2
Multimodal score67.664.8
Reasoning score66.868.9
Blended price $/M$3.38$3.00
Output speed209 tok/s56 tok/s
Time to first answer17.0 s12.7 s
Context window1M500k
SimpleBench70.0%
WeirdML46.4%
APEX-Agents34.2%
FrontierCode42.4%
Epoch Capabilities Index153.9
LMArena Text14761471
LMArena Hard Prompts14931495
LMArena Coding15031523
LMArena WebDev14911556
LMArena Vision13061291
LMArena Agent3.9
LiveBench75.8%
LiveBench Reasoning87.2%
LiveBench Coding68.6%
LiveBench Agentic Coding56.5%
LiveBench Mathematics90.8%
LiveBench Data Analysis73.0%
LiveBench Language82.8%
AA Intelligence Index33.639.1
IFBench74.6%
AA-LCR74.3%79.3%
MMMU-Pro83.9%80.4%
AA-Omniscience20.825.3
Terminal-Bench Hard39.4%
GPQA Diamond (AA)92.1%93.1%
Humanity's Last Exam (AA)41.3%42.7%
SciCode (AA)55.0%
τ²-Bench Telecom (AA)95.6%
Kagi LLM Benchmark83.5%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.