BenchLeader

Claude Opus 5 vs Qwen3.8 2.4T A95B

Verdict
  • Claude Opus 5 (high) leads on quality: 70.2 vs 65.2.
  • Claude Opus 5 (high) is stronger in coding, composite, human preference, knowledge, maths, multimodal.
  • Qwen3.8 2.4T A95B is stronger in agents & tools, long context, reasoning.
  • Qwen3.8 2.4T A95B is 3.3× cheaper ($3.00 vs $10.00 per 1M blended).
  • Claude Opus 5 (high) streams 1.4× faster (54 vs 38 tokens per second).
MetricClaude Opus 5 (high)Qwen3.8 2.4T A95B
BenchLeader Index70.265.2
Agents & tools score66.078.3
Coding score71.9
Composite score90.279.8
Human preference score69.7
Knowledge score80.366.3
Long context score65.966.6
Maths score72.3
Multimodal score67.6
Reasoning score75.680.5
Blended price $/M$10.00$3.00
Output speed54 tok/s38 tok/s
Time to first answer16.7 s55.2 s
Context window1M984k
Terminal-Bench50.3%
OSWorld-Verified 2.029.0%
SciCode54.3%
WeirdML91.6%
LMArena Text1493
LMArena Hard Prompts1517
LMArena Coding1533
LMArena WebDev1660
LMArena Vision1321
LMArena Agent10.2
AA Intelligence Index48.240.0
AA-LCR79.0%80.3%
MMMU-Pro82.4%
AA-Omniscience33.74.3
GPQA Diamond (AA)93.7%93.5%
Humanity's Last Exam (AA)52.8%42.5%
SciCode (AA)55.4%54.0%
ARC-AGI-197.5%
ARC-AGI-288.3%
ARC-AGI-330.2%
CritPt28.3%20.0%
GDPval (AA)56.5%56.4%
τ²-Bench Banking (AA)44.7%49.1%
LMArena Maths1525
LMArena Creative Writing1474
LMArena Instruction Following1499
LMArena Multi-turn1484
LMArena Longer Queries1508
LMArena Document1490
DeepSWE72.8%
CursorBench66.7%
ALE-Bench2164.6

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

Claude Opus 5 vs Qwen3.8 2.4T A95B: questions

Is Claude Opus 5 better than Qwen3.8 2.4T A95B?
Claude Opus 5 (high) leads on quality: 70.2 vs 65.2. The BenchLeader Index combines every independent quality benchmark; Claude Opus 5 (high) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is Claude Opus 5 better than Qwen3.8 2.4T A95B for agentic tasks?
Qwen3.8 2.4T A95B scores higher in agentic tasks (78 vs 66 on the category index, where 50 is average).
Which is cheaper, Claude Opus 5 or Qwen3.8 2.4T A95B?
Qwen3.8 2.4T A95B is cheaper: $3.00 against $10.00 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Opus 5 or Qwen3.8 2.4T A95B?
Claude Opus 5 streams faster: 54 against 38 output tokens per second.
Which has the larger context window?
Claude Opus 5 accepts more context: 1M against 984k tokens.