BenchLeader

Gemini 3.5 Flash vs Qwen3.8 Max

Verdict
  • Qwen3.8 Max (max) leads on quality: 64.2 vs 63.0.
  • Gemini 3.5 Flash (medium) is stronger in agents & tools, instruction following, knowledge, maths, multimodal.
  • Qwen3.8 Max (max) is stronger in coding, composite, human preference, long context, reasoning.
  • They cost about the same ($3.00 per 1M blended).
  • Gemini 3.5 Flash (medium) streams 5.6× faster (206 vs 37 tokens per second).
MetricGemini 3.5 Flash (medium)Qwen3.8 Max (max)
BenchLeader Index63.064.2
Agents & tools score67.958.2
Coding score56.864.6
Composite score67.670.7
Human preference score67.168.0
Instruction following score72.7–
Knowledge score71.162.7
Long context score62.365.4
Maths score66.865.1
Multimodal score66.566.2
Reasoning score61.772.2
Blended price $/M$3.38$3.00
Output speed206 tok/s37 tok/s
Time to first answer13.6 s58.9 s
Context window1.0M1M
Terminal-Bench–27.0%
SciCode–53.2%
APEX-Agents–63.3%
ProofBench–58.0%
Epoch Capabilities Index–156.4
LMArena Text14761483
LMArena Hard Prompts14951504
LMArena Coding15061524
LMArena WebDev14911672
LMArena Vision13091314
LMArena Agent–2.3
LiveBench–78.5%
LiveBench Reasoning–88.2%
LiveBench Coding–72.9%
LiveBench Agentic Coding–64.7%
LiveBench Mathematics–91.3%
LiveBench Data Analysis–78.4%
LiveBench Language–79.7%
LiveBench Instruction Following–74.1%
AA Intelligence Index v4.3.233.645.4
IFBench74.6%–
AA-LCR74.3%80.3%
MMMU-Pro83.9%82.8%
AA-Omniscience20.812.0
Terminal-Bench Hard39.4%–
GPQA Diamond (AA)92.1%92.8%
Humanity's Last Exam (AA)41.3%43.1%
SciCode (AA)–53.2%
τ²-Bench Telecom (AA)95.6%–
LiveCodeBench–87.8%
MMLU-Pro–88.6%
IOI–68.9%
LegalBench–83.6%
CorpFin–65.8%
TaxEval–75.5%
Terminal-Bench 2.1 (Vals)–67.4%
SWE-bench (Vals)–85.6%
GPQA Diamond (Vals)–93.7%
Vals Index–48.3
CritPt10.9%20.0%
GDPval-AA v2.1–58.6%
τ³-Banking (AA)–51.3%
ITBench SRE (AA)–40.3%
Analyst Agent (AA)–45.0%
APEX-Agents (AA)–42.4%
Code Migration–24.0%
CyberBench–28.6%
Excel Modeling Benchmark–60.1%
Finance Agent v2–50.6%
Harvey's Legal Agent Benchmark–10.4%
Legal Research Bench–47.6%
MedCode–40.7%
MedScribe–85.0%
MMMU-Pro (Vals)–88.0%
MortgageTax–64.0%
MysteryMechanism–23.9%
ProgramBench–0.0%
Public Benefits Bench–67.1%
SAGE–51.3%
SkillsBench–42.0%
Tax Agent Bench–66.0%
Terminal-Bench 4.0 (Vals)–34.3%
Terminal-Bench Science–1.4%
Vals Multimodal Index–65.4%
Vibe Code Bench 1-100–12.8%
Vibe Code Bench v1.1–64.7%
LMArena Maths14821497
LMArena Creative Writing14691470
LMArena Instruction Following14661474
LMArena Multi-turn14781492
LMArena Longer Queries14801492
LMArena Document1461–
Surface Evolver Bench58.1%–
DeepSWE v1.137.4%–
GDP.pdf14.0%–
Terminal-Bench 4.0 (AA)–38.9%
Terminal-Bench 2.1 (AA)–88.8%
AutomationBench–56.2%
GDP.pdf–22.8%
MLCR–20.0%
EnterpriseOps-Gym–47.6%
AA-Omniscience: accuracy51.0%31.9%
AA-Omniscience: non-hallucination38.2%71.2%
AA-Briefcase v1.1–1617

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Gemini 3.5 Flash vs Qwen3.8 Max: questions

Is Gemini 3.5 Flash better than Qwen3.8 Max?
Qwen3.8 Max (max) leads on quality: 64.2 vs 63.0. The BenchLeader Index combines every independent quality benchmark; Qwen3.8 Max (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Gemini 3.5 Flash better than Qwen3.8 Max for coding?
Qwen3.8 Max scores higher in coding (65 vs 57 on the category index, where 50 is average).
Is Gemini 3.5 Flash better than Qwen3.8 Max for agentic tasks?
Gemini 3.5 Flash scores higher in agentic tasks (68 vs 58 on the category index, where 50 is average).
Which is cheaper, Gemini 3.5 Flash or Qwen3.8 Max?
Qwen3.8 Max is cheaper: $3.00 against $3.38 per million tokens, blended at three input tokens per output token.
Which is faster, Gemini 3.5 Flash or Qwen3.8 Max?
Gemini 3.5 Flash streams faster: 206 against 37 output tokens per second.
Which has the larger context window?
Gemini 3.5 Flash accepts more context: 1.0M against 1M tokens.