BenchLeader

Agnes 3.0 Flash vs Claude Sonnet 5

Verdict
  • Agnes 3.0 Flash and Claude Sonnet 5 (high) are level on quality (62.3 vs 62.2).
  • Agnes 3.0 Flash is stronger in agents & tools, composite, long context, reasoning.
  • Claude Sonnet 5 (high) is stronger in knowledge, coding, human preference, maths, multimodal.
  • Agnes 3.0 Flash is 53× cheaper ($0.075 vs $4.00 per 1M blended).
MetricAgnes 3.0 FlashClaude Sonnet 5 (high)
BenchLeader Index62.362.2
Agents & tools score76.862.0
Composite score74.069.6
Knowledge score59.162.4
Long context score67.064.7
Reasoning score71.169.6
Coding score61.5
Human preference score65.9
Maths score66.9
Multimodal score62.6
Blended price $/M$0.075$4.00
Output speed64 tok/s
Time to first answer2.9 s
Context window1M1M
SciCode48.6%
WeirdML68.8%
LMArena Text1461
LMArena Hard Prompts1489
LMArena Coding1521
LMArena WebDev1537
LMArena Vision1277
LMArena Agent6
AA Intelligence Index35.532.0
AA-LCR81.0%76.7%
AA-Omniscience-10.6-3.7
GPQA Diamond (AA)92.4%
Humanity's Last Exam (AA)38.5%35.7%
SciCode (AA)51.6%54.3%
CritPt15.1%15.1%
GDPval (AA)53.7%40.4%
τ²-Bench Banking (AA)47.6%
LMArena Maths1476
LMArena Creative Writing1437
LMArena Instruction Following1465
LMArena Multi-turn1473
LMArena Longer Queries1482
LMArena Document1476
DeepSWE48.2%
CursorBench56.9%
ALE-Bench1463.1

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

Agnes 3.0 Flash vs Claude Sonnet 5: questions

Is Agnes 3.0 Flash better than Claude Sonnet 5?
Agnes 3.0 Flash and Claude Sonnet 5 (high) are level on quality (62.3 vs 62.2). The BenchLeader Index combines every independent quality benchmark; Agnes 3.0 Flash is ahead overall as of 2026-09-19, but check the category scores for your use.
Is Agnes 3.0 Flash better than Claude Sonnet 5 for agentic tasks?
Agnes 3.0 Flash scores higher in agentic tasks (77 vs 62 on the category index, where 50 is average).
Which is cheaper, Agnes 3.0 Flash or Claude Sonnet 5?
Agnes 3.0 Flash is cheaper: $0.075 against $4.00 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Both accept 1M tokens of context.