BenchLeader

Gemini 3.8 Flash vs Grok 4.5

Verdict
  • Gemini 3.8 Flash (medium) and Grok 4.5 are level on quality (63.5 vs 63.4).
  • Gemini 3.8 Flash (medium) is stronger in coding, composite, knowledge, long context, multimodal.
  • Grok 4.5 is stronger in agents & tools, human preference, reasoning.
  • Gemini 3.8 Flash (medium) is 2.0× cheaper ($1.50 vs $3.00 per 1M blended).
  • Gemini 3.8 Flash (medium) streams 4.9× faster (273 vs 56 tokens per second).
MetricGemini 3.8 Flash (medium)Grok 4.5
BenchLeader Index63.563.4
Coding score63.659.9
Composite score79.466.0
Knowledge score77.676.0
Long context score68.766.2
Multimodal score68.964.8
Agents & tools score58.6
Human preference score66.9
Reasoning score68.9
Blended price $/M$1.50$3.00
Output speed273 tok/s56 tok/s
Time to first answer13.8 s12.7 s
Context window1M500k
SimpleBench70.0%
SciCode54.4%
WeirdML46.4%
APEX-Agents34.2%
FrontierCode42.4%
Epoch Capabilities Index153.9
LMArena Text1471
LMArena Hard Prompts1495
LMArena Coding1523
LMArena WebDev1556
LMArena Vision1291
LMArena Agent3.9
LiveBench75.8%
LiveBench Reasoning87.2%
LiveBench Coding68.6%
LiveBench Agentic Coding56.5%
LiveBench Mathematics90.8%
LiveBench Data Analysis73.0%
LiveBench Language82.8%
AA Intelligence Index4039.1
AA-LCR84.0%79.3%
MMMU-Pro84.2%80.4%
AA-Omniscience28.625.3
GPQA Diamond (AA)93.5%93.1%
Humanity's Last Exam (AA)42.1%42.7%
SciCode (AA)55.1%55.0%
Kagi LLM Benchmark83.5%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.