BenchLeader

Qwen3 8 vs Qwen3.8 2.4T A95B

Verdict
  • Qwen3 8 (max) leads on quality: 66.5 vs 65.2.
  • Qwen3 8 (max) is stronger in coding, human preference, maths, multimodal.
  • Qwen3.8 2.4T A95B is stronger in agents & tools, composite, knowledge, reasoning.
  • They cost about the same ($2.67 per 1M blended).
MetricQwen3 8 (max)Qwen3.8 2.4T A95B
BenchLeader Index66.565.2
Agents & tools score68.378.3
Coding score66.5
Composite score74.379.8
Human preference score68.2
Knowledge score63.566.3
Long context score66.666.6
Maths score67.0
Multimodal score67.4
Reasoning score76.380.5
Blended price $/M$2.67$3.00
Output speed39 tok/s38 tok/s
Time to first answer55.1 s55.2 s
Context window1M984k
SciCode52.9%
ProofBench58.0%
Epoch Capabilities Index156.6
LMArena Text1481
LMArena Hard Prompts1503
LMArena Coding1522
LMArena WebDev1671
LMArena Vision1315
LMArena Agent3.3
LiveBench78.5%
LiveBench Reasoning88.2%
LiveBench Coding72.9%
LiveBench Agentic Coding64.7%
LiveBench Mathematics91.3%
LiveBench Data Analysis78.4%
LiveBench Language79.7%
LiveBench Instruction Following74.1%
AA Intelligence Index45.440.0
AA-LCR80.3%80.3%
MMMU-Pro82.8%
AA-Omniscience12.04.3
GPQA Diamond (AA)92.8%93.5%
Humanity's Last Exam (AA)43.1%42.5%
SciCode (AA)53.2%54.0%
LiveCodeBench87.8%
MMLU-Pro88.6%
IOI68.9%
LegalBench83.6%
CorpFin65.8%
TaxEval75.5%
Terminal-Bench 2.1 (Vals)67.4%
SWE-bench (Vals)85.6%
GPQA Diamond (Vals)93.7%
Vals Index51.8
CritPt20.0%20.0%
GDPval (AA)58.2%56.4%
τ²-Bench Banking (AA)51.3%49.1%
Code Migration24.0%
Excel Modeling Benchmark60.1%
Finance Agent v250.6%
Harvey's Legal Agent Benchmark10.4%
Legal Research Bench47.6%
MedCode40.7%
MedScribe85.0%
MMMU-Pro (Vals)88.0%
MortgageTax64.0%
MysteryMechanism23.9%
ProgramBench0.0%
Public Benefits Bench67.1%
SAGE51.3%
SkillsBench42.0%
Tax Agent Bench66.0%
Terminal-Bench 4.0 (Vals)24.8%
Terminal-Bench Science1.4%
Vals Multimodal Index65.4%
Vibe Code Bench 1-10012.8%
Vibe Code Bench v1.164.7%
LMArena Maths1498
LMArena Creative Writing1472
LMArena Instruction Following1477
LMArena Multi-turn1495
LMArena Longer Queries1492

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

Qwen3 8 vs Qwen3.8 2.4T A95B: questions

Is Qwen3 8 better than Qwen3.8 2.4T A95B?
Qwen3 8 (max) leads on quality: 66.5 vs 65.2. The BenchLeader Index combines every independent quality benchmark; Qwen3 8 (max) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is Qwen3 8 better than Qwen3.8 2.4T A95B for agentic tasks?
Qwen3.8 2.4T A95B scores higher in agentic tasks (78 vs 68 on the category index, where 50 is average).
Which is cheaper, Qwen3 8 or Qwen3.8 2.4T A95B?
Qwen3 8 is cheaper: $2.67 against $3.00 per million tokens, blended at three input tokens per output token.
Which is faster, Qwen3 8 or Qwen3.8 2.4T A95B?
Qwen3 8 streams faster: 39 against 38 output tokens per second.
Which has the larger context window?
Qwen3 8 accepts more context: 1M against 984k tokens.