BenchLeader

Qwen3.8 2.4T A95B vs Qwen3.8-Flash-Next

Verdict
  • Qwen3.8 2.4T A95B leads on quality: 65.2 vs 61.6.
  • Qwen3.8 2.4T A95B is stronger in agents & tools, composite, knowledge, long context, reasoning.
  • Qwen3.8-Flash-Next is stronger in coding, multimodal.
  • Qwen3.8-Flash-Next is 13× cheaper ($0.230 vs $3.00 per 1M blended).
  • Qwen3.8-Flash-Next streams 1.5× faster (55 vs 38 tokens per second).
MetricQwen3.8 2.4T A95BQwen3.8-Flash-Next
BenchLeader Index65.261.6
Agents & tools score78.365.8
Composite score79.867.4
Knowledge score66.359.6
Long context score66.666.3
Reasoning score80.563.4
Coding score71.2
Multimodal score64.3
Blended price $/M$3.00$0.230
Output speed38 tok/s55 tok/s
Time to first answer55.2 s39.0 s
Context window984k256k
LMArena WebDev1635
LMArena Agent-0.1
LiveBench76.2%
LiveBench Reasoning87.4%
LiveBench Coding72.5%
LiveBench Agentic Coding61.6%
LiveBench Mathematics85.8%
LiveBench Data Analysis74.2%
LiveBench Language74.6%
LiveBench Instruction Following77.1%
AA Intelligence Index40.039.9
AA-LCR80.3%79.7%
MMMU-Pro79.8%
AA-Omniscience4.3-9.7
GPQA Diamond (AA)93.5%92.3%
Humanity's Last Exam (AA)42.5%38.0%
SciCode (AA)54.0%50.6%
CritPt20.0%11.1%
GDPval (AA)56.4%57.4%
τ²-Bench Banking (AA)49.1%45.4%

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

Qwen3.8 2.4T A95B vs Qwen3.8-Flash-Next: questions

Is Qwen3.8 2.4T A95B better than Qwen3.8-Flash-Next?
Qwen3.8 2.4T A95B leads on quality: 65.2 vs 61.6. The BenchLeader Index combines every independent quality benchmark; Qwen3.8 2.4T A95B is ahead overall as of 2026-09-19, but check the category scores for your use.
Is Qwen3.8 2.4T A95B better than Qwen3.8-Flash-Next for agentic tasks?
Qwen3.8 2.4T A95B scores higher in agentic tasks (78 vs 66 on the category index, where 50 is average).
Which is cheaper, Qwen3.8 2.4T A95B or Qwen3.8-Flash-Next?
Qwen3.8-Flash-Next is cheaper: $0.230 against $3.00 per million tokens, blended at three input tokens per output token.
Which is faster, Qwen3.8 2.4T A95B or Qwen3.8-Flash-Next?
Qwen3.8-Flash-Next streams faster: 55 against 38 output tokens per second.
Which has the larger context window?
Qwen3.8 2.4T A95B accepts more context: 984k against 256k tokens.