BenchLeader

Gemini 4 Argon vs Qwen3.6 Max

Verdict
  • Gemini 4 Argon (high) leads on quality: 70.0 vs 62.5.
  • Gemini 4 Argon (high) is stronger in coding, composite, human preference, knowledge, maths, reasoning.
  • Qwen3.6 Max (max) is stronger in agents & tools, long context, instruction following.
  • Qwen3.6 Max (max) is 1.4× cheaper ($2.92 vs $4.00 per 1M blended).
MetricGemini 4 Argon (high)Qwen3.6 Max (max)
BenchLeader Index70.062.5
Agents & tools score67.571.9
Coding score72.156.9
Composite score89.261.6
Human preference score73.265.2
Knowledge score75.762.6
Long context score65.065.6
Maths score71.965.3
Reasoning score82.058.4
Instruction following score–74.4
Blended price $/M$4.00$2.92
Output speed–71 tok/s
Time to first answer–31.3 s
Context window1M262k
GPQA Diamond–87.4%
OTIS Mock AIME–91.1%
SWE-bench Verified (Epoch)–76.7%
SimpleQA Verified–52.0%
SimpleBench–63.0%
SciCode61.8%–
Epoch Capabilities Index–149.2
LMArena Text15251460
LMArena Hard Prompts15511482
LMArena Coding15621511
LMArena WebDev16781482
LMArena Agent9.3–
AA Intelligence Index v4.3.252.628.4
IFBench–76.6%
AA-LCR79.7%80.7%
AA-Omniscience42.49.2
Terminal-Bench Hard–43.9%
GPQA Diamond (AA)–88.8%
Humanity's Last Exam (AA)57.1%30.8%
SciCode (AA)61.8%–
τ²-Bench Telecom (AA)–95.9%
IOI100.0%–
LegalBench88.3%–
CorpFin–66.5%
SWE-bench (Vals)–72.8%
Vals Index68.9–
CritPt27.1%3.7%
GDPval-AA v2.156.3%–
BioMysteryBench76.3%–
CaseLaw v2–47.9%
Code Migration68.2%–
CUA-bench4.8%–
CyberBench77.9%–
Excel Modeling Benchmark75.2%–
Finance Agent v265.4%–
Harvey's Legal Agent Benchmark19.6%–
Legal Research Bench54.8%–
MedCode58.8%–
MedScribe87.4%–
MysteryMechanism45.5%–
ProgramBench2.5%–
Public Benefits Bench69.8%–
SAGE53.6%–
SREBench44.3%–
Tax Agent Bench76.2%–
Terminal-Bench 2.0 (Vals)–51.7%
Terminal-Bench 4.0 (Vals)57.6%–
Terminal-Bench Science44.3%–
Vibe Code Bench v1.191.9%–
LMArena Maths15281476
LMArena Creative Writing15191437
LMArena Instruction Following15291451
LMArena Multi-turn15531469
LMArena Longer Queries15451473
Chess Puzzles–20.0%
LMCA–42.5%
DTBench–87.2%
Vending-Bench 2–4254.2
FrontierSWE55.0%–
Terminal-Bench 4.0 (AA)57.1%–
AutomationBench77.5%–
GDP.pdf21.8%–
AA-Omniscience: accuracy49.9%37.9%
AA-Omniscience: non-hallucination84.9%53.8%
AA-Briefcase v1.11488–

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Gemini 4 Argon vs Qwen3.6 Max: questions

Is Gemini 4 Argon better than Qwen3.6 Max?
Gemini 4 Argon (high) leads on quality: 70.0 vs 62.5. The BenchLeader Index combines every independent quality benchmark; Gemini 4 Argon (high) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Gemini 4 Argon better than Qwen3.6 Max for coding?
Gemini 4 Argon scores higher in coding (72 vs 57 on the category index, where 50 is average).
Is Gemini 4 Argon better than Qwen3.6 Max for agentic tasks?
Qwen3.6 Max scores higher in agentic tasks (72 vs 68 on the category index, where 50 is average).
Which is cheaper, Gemini 4 Argon or Qwen3.6 Max?
Qwen3.6 Max is cheaper: $2.92 against $4.00 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Gemini 4 Argon accepts more context: 1M against 262k tokens.