BenchLeader

Qwen3 8 vs Qwen3.8 27B

Verdict
  • Qwen3 8 (max) leads on quality: 66.5 vs 59.2.
  • Qwen3 8 (max) is stronger in agents & tools, coding, composite, human preference, knowledge, maths, multimodal, reasoning.
  • Qwen3.8 27B (xhigh) is stronger in long context.
  • Qwen3.8 27B (xhigh) is 2.8× cheaper ($0.938 vs $2.67 per 1M blended).
MetricQwen3 8 (max)Qwen3.8 27B (xhigh)
BenchLeader Index66.559.2
Agents & tools score68.367.3
Coding score66.553.2
Composite score74.372.0
Human preference score68.2
Knowledge score63.556.4
Long context score66.667.5
Maths score67.0
Multimodal score67.460.7
Reasoning score76.352.5
Blended price $/M$2.67$0.938
Output speed39 tok/s42 tok/s
Time to first answer55.1 s51.0 s
Context window1M262k
SciCode52.9%44.7%
ProofBench58.0%
Epoch Capabilities Index156.6
LMArena Text1481
LMArena Hard Prompts1503
LMArena Coding1522
LMArena WebDev1671
LMArena Vision1315
LMArena Agent3.3
LiveBench78.5%
LiveBench Reasoning88.2%
LiveBench Coding72.9%
LiveBench Agentic Coding64.7%
LiveBench Mathematics91.3%
LiveBench Data Analysis78.4%
LiveBench Language79.7%
LiveBench Instruction Following74.1%
AA Intelligence Index45.433.9
AA-LCR80.3%82.0%
MMMU-Pro82.8%76.3%
AA-Omniscience12.0-10.0
GPQA Diamond (AA)92.8%90.5%
Humanity's Last Exam (AA)43.1%33.9%
SciCode (AA)53.2%46.6%
LiveCodeBench87.8%84.0%
MMLU-Pro88.6%84.3%
IOI68.9%39.1%
LegalBench83.6%82.4%
CorpFin65.8%
TaxEval75.5%70.8%
Terminal-Bench 2.1 (Vals)67.4%58.4%
SWE-bench (Vals)85.6%86.0%
GPQA Diamond (Vals)93.7%88.9%
Vals Index51.848.5
CritPt20.0%5.4%
GDPval (AA)58.2%48.2%
τ²-Bench Banking (AA)51.3%48.0%
Code Migration24.0%14.2%
Excel Modeling Benchmark60.1%59.7%
Finance Agent v250.6%48.5%
Harvey's Legal Agent Benchmark10.4%11.3%
Legal Research Bench47.6%36.1%
MedCode40.7%28.7%
MedScribe85.0%83.8%
MMMU-Pro (Vals)88.0%83.9%
MortgageTax64.0%64.9%
MysteryMechanism23.9%
ProgramBench0.0%0.0%
Public Benefits Bench67.1%
SAGE51.3%52.4%
SkillsBench42.0%38.1%
Tax Agent Bench66.0%
Terminal-Bench 4.0 (Vals)24.8%4.0%
Terminal-Bench Science1.4%1.4%
Vals Multimodal Index65.4%
Vibe Code Bench 1-10012.8%
Vibe Code Bench v1.164.7%64.8%
LMArena Maths1498
LMArena Creative Writing1472
LMArena Instruction Following1477
LMArena Multi-turn1495
LMArena Longer Queries1492
Surface Evolver Bench45.0%

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

Qwen3 8 vs Qwen3.8 27B: questions

Is Qwen3 8 better than Qwen3.8 27B?
Qwen3 8 (max) leads on quality: 66.5 vs 59.2. The BenchLeader Index combines every independent quality benchmark; Qwen3 8 (max) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is Qwen3 8 better than Qwen3.8 27B for coding?
Qwen3 8 scores higher in coding (67 vs 53 on the category index, where 50 is average).
Is Qwen3 8 better than Qwen3.8 27B for agentic tasks?
Qwen3 8 scores higher in agentic tasks (68 vs 67 on the category index, where 50 is average).
Which is cheaper, Qwen3 8 or Qwen3.8 27B?
Qwen3.8 27B is cheaper: $0.938 against $2.67 per million tokens, blended at three input tokens per output token.
Which is faster, Qwen3 8 or Qwen3.8 27B?
Qwen3.8 27B streams faster: 42 against 39 output tokens per second.
Which has the larger context window?
Qwen3 8 accepts more context: 1M against 262k tokens.