BenchLeader

Cybench

Unguided capture-the-flag cybersecurity tasks.

Published by
Cybenchdata via Epoch AI Benchmarking Hub
Category
Agents & tools
Index weight
0.5
Models
21
Data as of
9 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

21 of 21
#
1Claude Opus 4.6Anthropic93.0%63.1
2Claude Opus 4.5Anthropic82.0%58.0
3Claude Sonnet 4.5Anthropic60.0%52.3
4Grok 4xAI43.0%070957.2
5Claude Opus 4.1Anthropic42.0%52.6
6Grok 4.1xAI39.0%48.8
7Claude Opus 4Anthropic38.0%50.7
8Claude Sonnet 4Anthropic35.0%49.5
9Grok 4 FastxAI30.0%46.5
10o3-miniOpenAI22.5%45.6
11Claude 3.7 SonnetAnthropic20.0%49.8
12GPT-4.5OpenAI17.5%49.6
13Claude 3.5 SonnetAnthropic17.5%45.1
14GPT-4oOpenAI12.5%42.3
15o1OpenAI10.0%53.2
16o1-miniOpenAI10.0%44.4
17Claude 3 OpusAnthropic10.0%38.9
18Llama 3.1 405BMetaopen7.5%42.6
19Mixtral 8x22BMistral AIopen7.5%39.0
20Gemini 1.5 Pro 001 Feb24Google7.5%
21Llama 3-70BMetaopen5.0%38.6