BenchLeader

Qwen3.6 Max vs Qwen3.8 Max

Verdict
  • Qwen3.8 Max (max) leads on quality: 64.2 vs 62.5.
  • Qwen3.6 Max (max) is stronger in agents & tools, instruction following, long context, maths.
  • Qwen3.8 Max (max) is stronger in coding, composite, human preference, knowledge, reasoning, multimodal.
  • They cost about the same ($2.92 per 1M blended).
  • Qwen3.6 Max (max) streams 1.9× faster (71 vs 37 tokens per second).
MetricQwen3.6 Max (max)Qwen3.8 Max (max)
BenchLeader Index62.564.2
Agents & tools score71.958.2
Coding score56.964.6
Composite score61.670.7
Human preference score65.268.0
Instruction following score74.4–
Knowledge score62.662.7
Long context score65.665.4
Maths score65.365.1
Reasoning score58.472.2
Multimodal score–66.2
Blended price $/M$2.92$3.00
Output speed71 tok/s37 tok/s
Time to first answer31.3 s58.9 s
Context window262k1M
GPQA Diamond87.4%–
OTIS Mock AIME91.1%–
SWE-bench Verified (Epoch)76.7%–
SimpleQA Verified52.0%–
Terminal-Bench–27.0%
SimpleBench63.0%–
SciCode–53.2%
APEX-Agents–63.3%
ProofBench–58.0%
Epoch Capabilities Index149.2156.4
LMArena Text14601483
LMArena Hard Prompts14821504
LMArena Coding15111524
LMArena WebDev14821672
LMArena Vision–1314
LMArena Agent–2.3
LiveBench–78.5%
LiveBench Reasoning–88.2%
LiveBench Coding–72.9%
LiveBench Agentic Coding–64.7%
LiveBench Mathematics–91.3%
LiveBench Data Analysis–78.4%
LiveBench Language–79.7%
LiveBench Instruction Following–74.1%
AA Intelligence Index v4.3.228.445.4
IFBench76.6%–
AA-LCR80.7%80.3%
MMMU-Pro–82.8%
AA-Omniscience9.212.0
Terminal-Bench Hard43.9%–
GPQA Diamond (AA)88.8%92.8%
Humanity's Last Exam (AA)30.8%43.1%
SciCode (AA)–53.2%
τ²-Bench Telecom (AA)95.9%–
LiveCodeBench–87.8%
MMLU-Pro–88.6%
IOI–68.9%
LegalBench–83.6%
CorpFin66.5%65.8%
TaxEval–75.5%
Terminal-Bench 2.1 (Vals)–67.4%
SWE-bench (Vals)72.8%85.6%
GPQA Diamond (Vals)–93.7%
Vals Index–48.3
CritPt3.7%20.0%
GDPval-AA v2.1–58.6%
τ³-Banking (AA)–51.3%
ITBench SRE (AA)–40.3%
Analyst Agent (AA)–45.0%
APEX-Agents (AA)–42.4%
CaseLaw v247.9%–
Code Migration–24.0%
CyberBench–28.6%
Excel Modeling Benchmark–60.1%
Finance Agent v2–50.6%
Harvey's Legal Agent Benchmark–10.4%
Legal Research Bench–47.6%
MedCode–40.7%
MedScribe–85.0%
MMMU-Pro (Vals)–88.0%
MortgageTax–64.0%
MysteryMechanism–23.9%
ProgramBench–0.0%
Public Benefits Bench–67.1%
SAGE–51.3%
SkillsBench–42.0%
Tax Agent Bench–66.0%
Terminal-Bench 2.0 (Vals)51.7%–
Terminal-Bench 4.0 (Vals)–34.3%
Terminal-Bench Science–1.4%
Vals Multimodal Index–65.4%
Vibe Code Bench 1-100–12.8%
Vibe Code Bench v1.1–64.7%
LMArena Maths14761497
LMArena Creative Writing14371470
LMArena Instruction Following14511474
LMArena Multi-turn14691492
LMArena Longer Queries14731492
Chess Puzzles20.0%–
LMCA42.5%–
DTBench87.2%–
Vending-Bench 24254.2–
Terminal-Bench 4.0 (AA)–38.9%
Terminal-Bench 2.1 (AA)–88.8%
AutomationBench–56.2%
GDP.pdf–22.8%
MLCR–20.0%
EnterpriseOps-Gym–47.6%
AA-Omniscience: accuracy37.9%31.9%
AA-Omniscience: non-hallucination53.8%71.2%
AA-Briefcase v1.1–1617

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Qwen3.6 Max vs Qwen3.8 Max: questions

Is Qwen3.6 Max better than Qwen3.8 Max?
Qwen3.8 Max (max) leads on quality: 64.2 vs 62.5. The BenchLeader Index combines every independent quality benchmark; Qwen3.8 Max (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Qwen3.6 Max better than Qwen3.8 Max for coding?
Qwen3.8 Max scores higher in coding (65 vs 57 on the category index, where 50 is average).
Is Qwen3.6 Max better than Qwen3.8 Max for agentic tasks?
Qwen3.6 Max scores higher in agentic tasks (72 vs 58 on the category index, where 50 is average).
Which is cheaper, Qwen3.6 Max or Qwen3.8 Max?
Qwen3.6 Max is cheaper: $2.92 against $3.00 per million tokens, blended at three input tokens per output token.
Which is faster, Qwen3.6 Max or Qwen3.8 Max?
Qwen3.6 Max streams faster: 71 against 37 output tokens per second.
Which has the larger context window?
Qwen3.8 Max accepts more context: 1M against 262k tokens.