BenchLeader

Deepseek v4 Flash Vision vs Qwen3 8

Verdict
  • Qwen3 8 (max) leads on quality: 66.5 vs 59.8.
  • Deepseek v4 Flash Vision (max) is stronger in agents & tools, long context.
  • Qwen3 8 (max) is stronger in composite, knowledge, multimodal, reasoning, coding, human preference, maths.
  • Deepseek v4 Flash Vision (max) is 10× cheaper ($0.262 vs $2.67 per 1M blended).
  • Deepseek v4 Flash Vision (max) streams 5.6× faster (215 vs 39 tokens per second).
MetricDeepseek v4 Flash Vision (max)Qwen3 8 (max)
BenchLeader Index59.866.5
Agents & tools score69.968.3
Composite score73.474.3
Knowledge score55.863.5
Long context score67.166.6
Multimodal score59.267.4
Reasoning score62.976.3
Coding score66.5
Human preference score68.2
Maths score67.0
Blended price $/M$0.262$2.67
Output speed215 tok/s39 tok/s
Time to first answer10.4 s55.1 s
Context window1M1M
SciCode52.9%
ProofBench58.0%
Epoch Capabilities Index156.6
LMArena Text1481
LMArena Hard Prompts1503
LMArena Coding1522
LMArena WebDev1671
LMArena Vision1315
LMArena Agent3.3
LiveBench78.5%
LiveBench Reasoning88.2%
LiveBench Coding72.9%
LiveBench Agentic Coding64.7%
LiveBench Mathematics91.3%
LiveBench Data Analysis78.4%
LiveBench Language79.7%
LiveBench Instruction Following74.1%
AA Intelligence Index35.045.4
AA-LCR81.3%80.3%
MMMU-Pro74.8%82.8%
AA-Omniscience-17.612.0
GPQA Diamond (AA)91.3%92.8%
Humanity's Last Exam (AA)34.5%43.1%
SciCode (AA)49.6%53.2%
LiveCodeBench87.8%
MMLU-Pro88.6%
IOI68.9%
LegalBench83.6%
CorpFin65.8%
TaxEval75.5%
Terminal-Bench 2.1 (Vals)67.4%
SWE-bench (Vals)85.6%
GPQA Diamond (Vals)93.7%
Vals Index51.8
CritPt10.9%20.0%
GDPval (AA)53.9%58.2%
τ²-Bench Banking (AA)41.0%51.3%
Code Migration24.0%
Excel Modeling Benchmark60.1%
Finance Agent v250.6%
Harvey's Legal Agent Benchmark10.4%
Legal Research Bench47.6%
MedCode40.7%
MedScribe85.0%
MMMU-Pro (Vals)88.0%
MortgageTax64.0%
MysteryMechanism23.9%
ProgramBench0.0%
Public Benefits Bench67.1%
SAGE51.3%
SkillsBench42.0%
Tax Agent Bench66.0%
Terminal-Bench 4.0 (Vals)24.8%
Terminal-Bench Science1.4%
Vals Multimodal Index65.4%
Vibe Code Bench 1-10012.8%
Vibe Code Bench v1.164.7%
LMArena Maths1498
LMArena Creative Writing1472
LMArena Instruction Following1477
LMArena Multi-turn1495
LMArena Longer Queries1492

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

Deepseek v4 Flash Vision vs Qwen3 8: questions

Is Deepseek v4 Flash Vision better than Qwen3 8?
Qwen3 8 (max) leads on quality: 66.5 vs 59.8. The BenchLeader Index combines every independent quality benchmark; Qwen3 8 (max) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is Deepseek v4 Flash Vision better than Qwen3 8 for agentic tasks?
Deepseek v4 Flash Vision scores higher in agentic tasks (70 vs 68 on the category index, where 50 is average).
Which is cheaper, Deepseek v4 Flash Vision or Qwen3 8?
Deepseek v4 Flash Vision is cheaper: $0.262 against $2.67 per million tokens, blended at three input tokens per output token.
Which is faster, Deepseek v4 Flash Vision or Qwen3 8?
Deepseek v4 Flash Vision streams faster: 215 against 39 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.