BenchLeader

Gemini 3.5 Flash vs Grok 4.3

Verdict
  • Gemini 3.5 Flash (medium) leads on quality: 64.1 vs 60.7.
  • Gemini 3.5 Flash (medium) is stronger in agents & tools, coding, composite, human preference, knowledge, multimodal, reasoning.
  • Grok 4.3 (medium) is stronger in instruction following, long context.
  • Grok 4.3 (medium) is 2.2× cheaper ($1.56 vs $3.38 per 1M blended).
  • Gemini 3.5 Flash (medium) streams 1.9× faster (217 vs 112 tokens per second).
MetricGemini 3.5 Flash (medium)Grok 4.3 (medium)
BenchLeader Index64.160.7
Agents & tools score68.060.1
Coding score58.9
Composite score71.560.3
Human preference score67.8
Instruction following score72.580.0
Knowledge score74.072.1
Long context score63.664.0
Multimodal score67.660.2
Reasoning score67.1
Blended price $/M$3.38$1.56
Output speed217 tok/s112 tok/s
Time to first answer15.5 s12.1 s
Context window1M1M
LMArena Text1476
LMArena Hard Prompts1493
LMArena Coding1503
LMArena WebDev1491
LMArena Vision1306
AA Intelligence Index33.624.8
IFBench74.6%83.3%
AA-LCR74.3%75.0%
MMMU-Pro83.9%75.8%
AA-Omniscience20.816.7
Terminal-Bench Hard39.4%30.3%
GPQA Diamond (AA)92.1%89.0%
Humanity's Last Exam (AA)41.3%30.0%
τ²-Bench Telecom (AA)95.6%91.2%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Gemini 3.5 Flash vs Grok 4.3: questions

Is Gemini 3.5 Flash better than Grok 4.3?
Gemini 3.5 Flash (medium) leads on quality: 64.1 vs 60.7. The BenchLeader Index combines every independent quality benchmark; Gemini 3.5 Flash (medium) is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Gemini 3.5 Flash better than Grok 4.3 for agentic tasks?
Gemini 3.5 Flash scores higher in agentic tasks (68 vs 60 on the category index, where 50 is average).
Which is cheaper, Gemini 3.5 Flash or Grok 4.3?
Grok 4.3 is cheaper: $1.56 against $3.38 per million tokens, blended at three input tokens per output token.
Which is faster, Gemini 3.5 Flash or Grok 4.3?
Gemini 3.5 Flash streams faster: 217 against 112 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.