ExploitBench
Developing working exploits for known vulnerabilities.
As of 19 Sept 2026, Claude Mythos Preview leads ExploitBench on BenchLeader with 73.8%, ahead of GPT-5.5 at 47.4%, across 9 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 9
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Given a vulnerable target, the agent must produce a working exploit, scored on capability.
How it is scored
Mean capability score, as published.
What to keep in mind
A dual-use capability measure; small model set.
- 1Claude Mythos Preview73.8%
- 2GPT-5.547.4%
- 3Claude Opus 4.726.5%
- 4Gemini 3.1 Pro26.1%
- 5Claude Sonnet 4.623.6%
- 6Kimi K2.618.4%
- 7GLM-5.118.1%
- 8Claude Haiku 4.513.7%
- 9MiniMax-M2.713.3%
9 of 9
| # | ||||
|---|---|---|---|---|
| 1 | 73.8% | – | 2026-04-07 | |
| 2 | 47.4% | 63.2 | 2026-04-23 | |
| 3 | 26.5% | 64.5 | 2026-04-16 | |
| 4 | 26.1% | 63.9 | 2026-02-19 | |
| 5 | 23.6% | 59.0 | 2026-02-17 | |
| 6 | 18.4% | 60.5 | 2026-04-20 | |
| 7 | 18.1% | 56.9 | 2026-04-07 | |
| 8 | 13.7% | 48.1 | 2025-10-15 | |
| 9 | 13.3% | 56.1 | 2026-03-18 |
Cite as: BenchLeader, “ExploitBench leaderboard”, https://www.benchleader.com/benchmarks/exploitbench, data as of 19 Sept 2026.
ExploitBench: questions
- What does ExploitBench measure?
- Given a vulnerable target, the agent must produce a working exploit, scored on capability. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads ExploitBench?
- Claude Mythos Preview leads ExploitBench with 73.8% as of 19 Sept 2026, ahead of GPT-5.5 at 47.4%.
- How many models have ExploitBench results?
- 9 model configurations have a ExploitBench result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs ExploitBench and how often is it updated?
- ExploitBench is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does ExploitBench count toward the BenchLeader Index?
- No. ExploitBench is shown for reference but left out of the composite index.