BenchLeader

Gemini 3.7 Flash vs Grok 4.5

Verdict
  • Gemini 3.7 Flash (medium) leads on quality: 64.4 vs 63.4.
  • Gemini 3.7 Flash (medium) is stronger in coding, composite, long context, multimodal.
  • Grok 4.5 is stronger in knowledge, reasoning, agents & tools, human preference.
  • Gemini 3.7 Flash (medium) is 2.0× cheaper ($1.50 vs $3.00 per 1M blended).
  • Gemini 3.7 Flash (medium) streams 5.1× faster (282 vs 56 tokens per second).
MetricGemini 3.7 Flash (medium)Grok 4.5
BenchLeader Index64.463.4
Coding score68.359.9
Composite score78.966.0
Knowledge score75.276.0
Long context score68.166.2
Multimodal score69.464.8
Reasoning score63.568.9
Agents & tools score58.6
Human preference score66.9
Blended price $/M$1.50$3.00
Output speed282 tok/s56 tok/s
Time to first answer5.3 s12.7 s
Context window1M500k
SimpleBench70.0%
SciCode57.9%
WeirdML46.4%
APEX-Agents34.2%
FrontierCode42.4%
Epoch Capabilities Index153.9
LMArena Text1471
LMArena Hard Prompts1495
LMArena Coding1523
LMArena WebDev1556
LMArena Vision1291
LMArena Agent3.9
LiveBench75.8%
LiveBench Reasoning87.2%
LiveBench Coding68.6%
LiveBench Agentic Coding56.5%
LiveBench Mathematics90.8%
LiveBench Data Analysis73.0%
LiveBench Language82.8%
AA Intelligence Index39.639.1
AA-LCR83.0%79.3%
MMMU-Pro84.7%80.4%
AA-Omniscience23.725.3
GPQA Diamond (AA)92.1%93.1%
Humanity's Last Exam (AA)39.0%42.7%
SciCode (AA)59.8%55.0%
Kagi LLM Benchmark83.5%
ARC-AGI-191.2%
ARC-AGI-263.8%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.