BenchLeader

Gemini 3.5 Flash vs Grok 4

Verdict
  • Gemini 3.5 Flash (medium) leads on quality: 62.2 vs 59.8.
  • Gemini 3.5 Flash (medium) is stronger in agents & tools, composite, human preference, instruction following, knowledge, multimodal, reasoning.
  • Grok 4 is stronger in coding, long context, maths.
  • Grok 4 is 2.2× cheaper ($1.56 vs $3.38 per 1M blended).
MetricGemini 3.5 Flash (medium)Grok 4
BenchLeader Index62.259.8
Agents & tools score63.353.4
Coding score58.761.2
Composite score65.3
Human preference score67.559.8
Instruction following score69.467.7
Knowledge score70.756.0
Long context score60.777.7
Multimodal score67.354.5
Reasoning score67.058.3
Maths score58.0
Blended price $/M$3.38$1.56
Output speed234 tok/s
Time to first answer11.2 s
Context window1.0M256k
GPQA Diamond87.0%
OTIS Mock AIME84.0%
Terminal-Bench27.2%
SimpleBench60.5%
Fiction.LiveBench 120k96.9%
Cybench43.0%
WeirdML45.7%
APEX-Agents15.2%
Epoch Capabilities Index146.4
LMArena Text14741411
LMArena Hard Prompts14931420
LMArena Coding15021435
LMArena WebDev1492
LMArena Vision13061210
AA Intelligence Index33.6
IFBench74.6%
AA-LCR74.3%
MMMU-Pro83.9%
AA-Omniscience20.8
Terminal-Bench Hard39.4%
GPQA Diamond (AA)92.1%
Humanity's Last Exam (AA)41.3%
τ²-Bench Telecom (AA)95.6%
AIME (Vals)90.6%
LiveCodeBench83.3%
MMLU-Pro85.3%
LegalBench83.2%
CorpFin66.0%
TaxEval65.1%
MedQA92.5%
MGSM90.9%
SWE-bench (Vals)57.8%
GPQA Diamond (Vals)88.1%
IMO 202521.4%
MathArena Apex2.1%
Kagi LLM Benchmark73.6%
IFEval (HELM)94.9%
Omni-MATH (HELM)60.3%
WildBench (HELM)79.7%
MMLU-Pro (HELM)85.1%
GPQA Diamond (HELM)72.6%
HELM Capabilities mean78.5%
Aider Polyglot79.6%
ARC-AGI-179.6%
ARC-AGI-229.4%
BFCL Overall63.0%

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

Gemini 3.5 Flash vs Grok 4: questions

Is Gemini 3.5 Flash better than Grok 4?
Gemini 3.5 Flash (medium) leads on quality: 62.2 vs 59.8. The BenchLeader Index combines every independent quality benchmark; Gemini 3.5 Flash (medium) is ahead overall as of 2026-09-13, but check the category scores for your use.
Is Gemini 3.5 Flash better than Grok 4 for coding?
Grok 4 scores higher in coding (61 vs 59 on the category index, where 50 is average).
Is Gemini 3.5 Flash better than Grok 4 for agentic tasks?
Gemini 3.5 Flash scores higher in agentic tasks (63 vs 53 on the category index, where 50 is average).
Which is cheaper, Gemini 3.5 Flash or Grok 4?
Grok 4 is cheaper: $1.56 against $3.38 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Gemini 3.5 Flash accepts more context: 1.0M against 256k tokens.