BenchLeader

Gemini 3.1 Pro vs Grok 4

Verdict
  • Gemini 3.1 Pro leads on quality: 62.8 vs 59.8.
  • Gemini 3.1 Pro is stronger in agents & tools, composite, instruction following, knowledge, maths, multimodal, reasoning.
  • Grok 4 is stronger in coding, long context.
  • Grok 4 is 2.9× cheaper ($1.56 vs $4.50 per 1M blended).
MetricGemini 3.1 ProGrok 4
BenchLeader Index62.859.8
Agents & tools score62.653.4
Coding score59.861.2
Composite score61.5
Human preference score59.859.8
Instruction following score68.467.7
Knowledge score64.956.0
Long context score65.277.7
Maths score60.758.0
Multimodal score65.854.5
Reasoning score69.858.3
Blended price $/M$4.50$1.56
Output speed115 tok/s
Time to first answer23.4 s
Context window1.0M256k
GPQA Diamond94.1%87.0%
FrontierMath Tiers 1–359.6%
FrontierMath Tier 426.8%
OTIS Mock AIME95.6%84.0%
Humanity's Last Exam46.4%
Terminal-Bench80.2%27.2%
SimpleBench79.6%60.5%
Fiction.LiveBench 120k96.9%
SciCode58.9%
Cybench43.0%
WeirdML72.1%45.7%
APEX-Agents33.5%15.2%
ProofBench26.0%
GSO-Bench22.6%
Epoch Capabilities Index155146.4
LMArena Text14871411
LMArena Hard Prompts15081420
LMArena Coding15201435
LMArena WebDev1447
LMArena Vision12951210
LMArena Agent-5.6
AA Intelligence Index30.4
IFBench77.1%
AA-LCR82.0%
MMMU-Pro82.4%
AA-Omniscience31.9
Terminal-Bench Hard53.8%
GPQA Diamond (AA)94.1%
Humanity's Last Exam (AA)47.0%
SciCode (AA)58.7%
τ²-Bench Telecom (AA)95.6%
AIME (Vals)90.6%
LiveCodeBench83.3%
MMLU-Pro85.3%
LegalBench83.2%
CorpFin66.0%
TaxEval65.1%
MedQA92.5%
MGSM90.9%
SWE-bench (Vals)57.8%
GPQA Diamond (Vals)88.1%
AIME 202698.3%
HMMT February 202694.7%
IMO 202521.4%
MathArena Apex60.9%2.1%
MultiChallenge71.4%
PRBench Finance41.9%
PRBench Legal44.0%
MultiNRC64.7%
HiL-Bench35.3%
TutorBench53.0%
EQ-Bench 41142
Kagi LLM Benchmark73.6%
IFEval (HELM)94.9%
Omni-MATH (HELM)60.3%
WildBench (HELM)79.7%
MMLU-Pro (HELM)85.1%
GPQA Diamond (HELM)72.6%
HELM Capabilities mean78.5%
Aider Polyglot79.6%
ARC-AGI-198.0%79.6%
ARC-AGI-277.1%29.4%
ARC-AGI-30.4%
BFCL Overall63.0%

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

Gemini 3.1 Pro vs Grok 4: questions

Is Gemini 3.1 Pro better than Grok 4?
Gemini 3.1 Pro leads on quality: 62.8 vs 59.8. The BenchLeader Index combines every independent quality benchmark; Gemini 3.1 Pro is ahead overall as of 2026-09-13, but check the category scores for your use.
Is Gemini 3.1 Pro better than Grok 4 for coding?
Grok 4 scores higher in coding (61 vs 60 on the category index, where 50 is average).
Is Gemini 3.1 Pro better than Grok 4 for agentic tasks?
Gemini 3.1 Pro scores higher in agentic tasks (63 vs 53 on the category index, where 50 is average).
Which is cheaper, Gemini 3.1 Pro or Grok 4?
Grok 4 is cheaper: $1.56 against $4.50 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Gemini 3.1 Pro accepts more context: 1.0M against 256k tokens.