BenchLeader

Gemini 3.1 Pro vs Qwen3.8 Max

Verdict
  • Qwen3.8 Max (max) leads on quality: 64.2 vs 62.1.
  • Gemini 3.1 Pro is stronger in instruction following, knowledge, long context.
  • Qwen3.8 Max (max) is stronger in agents & tools, coding, composite, human preference, maths, multimodal, reasoning.
  • Qwen3.8 Max (max) is 1.5× cheaper ($3.00 vs $4.50 per 1M blended).
  • Gemini 3.1 Pro streams 3.1× faster (114 vs 37 tokens per second).
MetricGemini 3.1 ProQwen3.8 Max (max)
BenchLeader Index62.164.2
Agents & tools score56.758.2
Coding score57.364.6
Composite score63.270.7
Human preference score59.568.0
Instruction following score68.4–
Knowledge score64.562.7
Long context score66.365.4
Maths score61.065.1
Multimodal score64.966.2
Reasoning score68.672.2
Blended price $/M$4.50$3.00
Output speed114 tok/s37 tok/s
Time to first answer24.8 s58.9 s
Context window1.0M1M
GPQA Diamond94.1%–
FrontierMath Tiers 1–359.6%–
FrontierMath Tier 426.8%–
OTIS Mock AIME95.6%–
Humanity's Last Exam46.4%–
Terminal-Bench80.2%27.0%
SimpleBench79.6%–
SciCode58.9%53.2%
WeirdML72.1%–
APEX-Agents35.3%63.3%
ProofBench26.0%58.0%
GSO-Bench22.6%–
Epoch Capabilities Index154.8156.4
LMArena Text14871483
LMArena Hard Prompts15081504
LMArena Coding15221524
LMArena WebDev14471672
LMArena Vision12961314
LMArena Agent-7.82.3
LiveBench–78.5%
LiveBench Reasoning–88.2%
LiveBench Coding–72.9%
LiveBench Agentic Coding–64.7%
LiveBench Mathematics–91.3%
LiveBench Data Analysis–78.4%
LiveBench Language–79.7%
LiveBench Instruction Following–74.1%
AA Intelligence Index v4.3.229.745.4
IFBench77.1%–
AA-LCR82.0%80.3%
MMMU-Pro82.4%82.8%
AA-Omniscience31.912.0
Terminal-Bench Hard53.8%–
GPQA Diamond (AA)94.1%92.8%
Humanity's Last Exam (AA)47.0%43.1%
SciCode (AA)58.7%53.2%
τ²-Bench Telecom (AA)95.6%–
LiveCodeBench–87.8%
MMLU-Pro–88.6%
IOI–68.9%
LegalBench–83.6%
CorpFin–65.8%
TaxEval–75.5%
Terminal-Bench 2.1 (Vals)–67.4%
SWE-bench (Vals)–85.6%
GPQA Diamond (Vals)–93.7%
Vals Index–48.3
AIME 202698.3%–
HMMT February 202694.7%–
MathArena Apex60.9%–
MultiChallenge71.4%–
PRBench Finance41.9%–
PRBench Legal44.0%–
MultiNRC64.7%–
HiL-Bench35.3%–
TutorBench53.0%–
EQ-Bench 41142–
ARC-AGI-198.0%–
ARC-AGI-277.1%–
ARC-AGI-30.4%–
CritPt17.7%20.0%
GDPval-AA v2.114.7%58.6%
τ³-Banking (AA)21.4%51.3%
ITBench SRE (AA)30.3%40.3%
Analyst Agent (AA)41.3%45.0%
APEX-Agents (AA)32.0%42.4%
Code Migration–24.0%
CyberBench–28.6%
Excel Modeling Benchmark–60.1%
Finance Agent v2–50.6%
Harvey's Legal Agent Benchmark–10.4%
Legal Research Bench–47.6%
MedCode–40.7%
MedScribe–85.0%
MMMU-Pro (Vals)–88.0%
MortgageTax–64.0%
MysteryMechanism–23.9%
ProgramBench–0.0%
Public Benefits Bench–67.1%
SAGE–51.3%
SkillsBench–42.0%
Tax Agent Bench–66.0%
Terminal-Bench 4.0 (Vals)–34.3%
Terminal-Bench Science–1.4%
Vals Multimodal Index–65.4%
Vibe Code Bench 1-100–12.8%
Vibe Code Bench v1.1–64.7%
FORTRESS29.8%–
MASK42.4%–
SWE Atlas: Codebase QnA13.5%–
SWE Atlas: Refactoring33.8%–
SWE Atlas: Test Writing29.8%–
LMArena Maths14881497
LMArena Creative Writing14811470
LMArena Instruction Following14801474
LMArena Multi-turn14971492
LMArena Longer Queries15001492
LMArena Document1459–
Chess Puzzles55.0%–
EBR-bench14.3%–
BALROG57.0%–
PostTrainBench22.0%–
ExploitBench26.1%–
CL-bench20.8%–
CL-bench Life16.9%–
METR Time Horizons77.0%–
DeepSWE v1.111.7%–
ForecastBench59.0%–
GBAEval0.8%–
ALE-Bench1160.6–
AlgoTune2.0–
Vending-Bench 2911.2–
Blueprint-Bench 226.5%–
GDP.pdf17.0%–
Terminal-Bench 4.0 (AA)4.0%38.9%
Terminal-Bench 2.1 (AA)73.8%88.8%
AutomationBench35.4%56.2%
GDP.pdf17.8%22.8%
MLCR15.6%20.0%
Harvey LAB0.0%–
EnterpriseOps-Gym42.2%47.6%
AA-Omniscience: accuracy54.9%31.9%
AA-Omniscience: non-hallucination49.1%71.2%
AA-Briefcase v1.14561617
BrowseComp31.2%–
DeepSearchQA60.2%–
FACTS Search83.7%–

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Gemini 3.1 Pro vs Qwen3.8 Max: questions

Is Gemini 3.1 Pro better than Qwen3.8 Max?
Qwen3.8 Max (max) leads on quality: 64.2 vs 62.1. The BenchLeader Index combines every independent quality benchmark; Qwen3.8 Max (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Gemini 3.1 Pro better than Qwen3.8 Max for coding?
Qwen3.8 Max scores higher in coding (65 vs 57 on the category index, where 50 is average).
Is Gemini 3.1 Pro better than Qwen3.8 Max for agentic tasks?
Qwen3.8 Max scores higher in agentic tasks (58 vs 57 on the category index, where 50 is average).
Which is cheaper, Gemini 3.1 Pro or Qwen3.8 Max?
Qwen3.8 Max is cheaper: $3.00 against $4.50 per million tokens, blended at three input tokens per output token.
Which is faster, Gemini 3.1 Pro or Qwen3.8 Max?
Gemini 3.1 Pro streams faster: 114 against 37 output tokens per second.
Which has the larger context window?
Gemini 3.1 Pro accepts more context: 1.0M against 1M tokens.