BenchLeader

DeepSeek V4.1 Flash vs Qwen3.8 Max

Verdict
  • Qwen3.8 Max (max) leads on quality: 64.2 vs 61.8.
  • DeepSeek V4.1 Flash (max) is stronger in composite, long context.
  • Qwen3.8 Max (max) is stronger in agents & tools, coding, human preference, knowledge, maths, multimodal, reasoning.
  • DeepSeek V4.1 Flash (max) is 5.7× cheaper ($0.525 vs $3.00 per 1M blended).
  • DeepSeek V4.1 Flash (max) streams 5.9× faster (217 vs 37 tokens per second).
MetricDeepSeek V4.1 Flash (max)Qwen3.8 Max (max)
BenchLeader Index61.864.2
Agents & tools score55.558.2
Coding score64.064.6
Composite score71.270.7
Human preference score66.968.0
Knowledge score59.362.7
Long context score67.365.4
Maths score61.965.1
Multimodal score60.966.2
Reasoning score65.272.2
Blended price $/M$0.525$3.00
Output speed217 tok/s37 tok/s
Time to first answer10.4 s58.9 s
Context window1M1M
GPQA Diamond89.8%–
FrontierMath Tiers 1–367.4%–
FrontierMath Tier 426.8%–
OTIS Mock AIME98.3%–
Terminal-Bench–27.0%
SciCode51.9%53.2%
APEX-Agents–63.3%
ProofBench–58.0%
Epoch Capabilities Index–156.4
LMArena Text14751483
LMArena Hard Prompts14991504
LMArena Coding15281524
LMArena WebDev16191672
LMArena Vision12771314
LMArena Agent3.82.3
LiveBench81.1%78.5%
LiveBench Reasoning86.7%88.2%
LiveBench Coding80.0%72.9%
LiveBench Agentic Coding77.3%64.7%
LiveBench Mathematics93.3%91.3%
LiveBench Data Analysis79.3%78.4%
LiveBench Language81.2%79.7%
LiveBench Instruction Following70.0%74.1%
AA Intelligence Index v4.3.239.545.4
AA-LCR84.0%80.3%
MMMU-Pro77.0%82.8%
AA-Omniscience-5.312.0
GPQA Diamond (AA)–92.8%
Humanity's Last Exam (AA)39.3%43.1%
SciCode (AA)51.9%53.2%
LiveCodeBench–87.8%
MMLU-Pro–88.6%
IOI–68.9%
LegalBench–83.6%
CorpFin–65.8%
TaxEval–75.5%
Terminal-Bench 2.1 (Vals)–67.4%
SWE-bench (Vals)–85.6%
GPQA Diamond (Vals)–93.7%
Vals Index–48.3
ARC-AGI-194.5%–
ARC-AGI-272.9%–
CritPt14.3%20.0%
GDPval-AA v2.155.0%58.6%
τ³-Banking (AA)–51.3%
ITBench SRE (AA)46.9%40.3%
Analyst Agent (AA)–45.0%
APEX-Agents (AA)–42.4%
Code Migration–24.0%
CyberBench–28.6%
Excel Modeling Benchmark–60.1%
Finance Agent v2–50.6%
Harvey's Legal Agent Benchmark6.7%10.4%
Legal Research Bench–47.6%
MedCode–40.7%
MedScribe–85.0%
MMMU-Pro (Vals)–88.0%
MortgageTax–64.0%
MysteryMechanism–23.9%
ProgramBench–0.0%
Public Benefits Bench64.3%67.1%
SAGE–51.3%
SkillsBench–42.0%
SREBench0.8%–
Tax Agent Bench–66.0%
Terminal-Bench 4.0 (Vals)–34.3%
Terminal-Bench Science–1.4%
Vals Multimodal Index–65.4%
Vibe Code Bench 1-100–12.8%
Vibe Code Bench v1.1–64.7%
LMArena Maths14931497
LMArena Creative Writing14391470
LMArena Instruction Following14771474
LMArena Multi-turn14661492
LMArena Longer Queries14821492
Mystery Game Puzzles43.0%–
ALE-Bench1092.3–
GDP.pdf19.8%–
Terminal-Bench 4.0 (AA)26.8%38.9%
Terminal-Bench 2.1 (AA)–88.8%
AutomationBench68.9%56.2%
GDP.pdf12.8%22.8%
MLCR22.8%20.0%
Harvey LAB1.7%–
EnterpriseOps-Gym–47.6%
AA-Omniscience: accuracy46.4%31.9%
AA-Omniscience: non-hallucination3.5%71.2%
AA-Briefcase v1.114221617
AA Openness Index44.4–

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

DeepSeek V4.1 Flash vs Qwen3.8 Max: questions

Is DeepSeek V4.1 Flash better than Qwen3.8 Max?
Qwen3.8 Max (max) leads on quality: 64.2 vs 61.8. The BenchLeader Index combines every independent quality benchmark; Qwen3.8 Max (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is DeepSeek V4.1 Flash better than Qwen3.8 Max for coding?
Qwen3.8 Max scores higher in coding (65 vs 64 on the category index, where 50 is average).
Is DeepSeek V4.1 Flash better than Qwen3.8 Max for agentic tasks?
Qwen3.8 Max scores higher in agentic tasks (58 vs 56 on the category index, where 50 is average).
Which is cheaper, DeepSeek V4.1 Flash or Qwen3.8 Max?
DeepSeek V4.1 Flash is cheaper: $0.525 against $3.00 per million tokens, blended at three input tokens per output token.
Which is faster, DeepSeek V4.1 Flash or Qwen3.8 Max?
DeepSeek V4.1 Flash streams faster: 217 against 37 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.