BenchLeader

Agnes 3.0 Flash vs Qwen3.8 2.4T A95B

Verdict
  • Qwen3.8 2.4T A95B leads on quality: 65.2 vs 62.3.
  • Agnes 3.0 Flash is stronger in long context.
  • Qwen3.8 2.4T A95B is stronger in agents & tools, composite, knowledge, reasoning.
  • Agnes 3.0 Flash is 40× cheaper ($0.075 vs $3.00 per 1M blended).
MetricAgnes 3.0 FlashQwen3.8 2.4T A95B
BenchLeader Index62.365.2
Agents & tools score76.878.3
Composite score74.079.8
Knowledge score59.166.3
Long context score67.066.6
Reasoning score71.180.5
Blended price $/M$0.075$3.00
Output speed38 tok/s
Time to first answer55.2 s
Context window1M984k
AA Intelligence Index35.540.0
AA-LCR81.0%80.3%
AA-Omniscience-10.64.3
GPQA Diamond (AA)92.4%93.5%
Humanity's Last Exam (AA)38.5%42.5%
SciCode (AA)51.6%54.0%
CritPt15.1%20.0%
GDPval (AA)53.7%56.4%
τ²-Bench Banking (AA)47.6%49.1%

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

Agnes 3.0 Flash vs Qwen3.8 2.4T A95B: questions

Is Agnes 3.0 Flash better than Qwen3.8 2.4T A95B?
Qwen3.8 2.4T A95B leads on quality: 65.2 vs 62.3. The BenchLeader Index combines every independent quality benchmark; Qwen3.8 2.4T A95B is ahead overall as of 2026-09-19, but check the category scores for your use.
Is Agnes 3.0 Flash better than Qwen3.8 2.4T A95B for agentic tasks?
Qwen3.8 2.4T A95B scores higher in agentic tasks (78 vs 77 on the category index, where 50 is average).
Which is cheaper, Agnes 3.0 Flash or Qwen3.8 2.4T A95B?
Agnes 3.0 Flash is cheaper: $0.075 against $3.00 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Agnes 3.0 Flash accepts more context: 1M against 984k tokens.