BenchLeader

Agnes 3.0 Flash vs Qwen3.8-Flash-Next

Verdict
  • Agnes 3.0 Flash and Qwen3.8-Flash-Next are level on quality (62.3 vs 61.6).
  • Agnes 3.0 Flash is stronger in agents & tools, composite, long context, reasoning.
  • Qwen3.8-Flash-Next is stronger in knowledge, coding, multimodal.
  • Agnes 3.0 Flash is 3.1× cheaper ($0.075 vs $0.230 per 1M blended).
MetricAgnes 3.0 FlashQwen3.8-Flash-Next
BenchLeader Index62.361.6
Agents & tools score76.865.8
Composite score74.067.4
Knowledge score59.159.6
Long context score67.066.3
Reasoning score71.163.4
Coding score71.2
Multimodal score64.3
Blended price $/M$0.075$0.230
Output speed55 tok/s
Time to first answer39.0 s
Context window1M256k
LMArena WebDev1635
LMArena Agent-0.1
LiveBench76.2%
LiveBench Reasoning87.4%
LiveBench Coding72.5%
LiveBench Agentic Coding61.6%
LiveBench Mathematics85.8%
LiveBench Data Analysis74.2%
LiveBench Language74.6%
LiveBench Instruction Following77.1%
AA Intelligence Index35.539.9
AA-LCR81.0%79.7%
MMMU-Pro79.8%
AA-Omniscience-10.6-9.7
GPQA Diamond (AA)92.4%92.3%
Humanity's Last Exam (AA)38.5%38.0%
SciCode (AA)51.6%50.6%
CritPt15.1%11.1%
GDPval (AA)53.7%57.4%
τ²-Bench Banking (AA)47.6%45.4%

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

Agnes 3.0 Flash vs Qwen3.8-Flash-Next: questions

Is Agnes 3.0 Flash better than Qwen3.8-Flash-Next?
Agnes 3.0 Flash and Qwen3.8-Flash-Next are level on quality (62.3 vs 61.6). The BenchLeader Index combines every independent quality benchmark; Agnes 3.0 Flash is ahead overall as of 2026-09-19, but check the category scores for your use.
Is Agnes 3.0 Flash better than Qwen3.8-Flash-Next for agentic tasks?
Agnes 3.0 Flash scores higher in agentic tasks (77 vs 66 on the category index, where 50 is average).
Which is cheaper, Agnes 3.0 Flash or Qwen3.8-Flash-Next?
Agnes 3.0 Flash is cheaper: $0.075 against $0.230 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Agnes 3.0 Flash accepts more context: 1M against 256k tokens.