BenchLeader

ExploitBench

Developing working exploits for known vulnerabilities.

As of 19 Sept 2026, Claude Mythos Preview leads ExploitBench on BenchLeader with 73.8%, ahead of GPT-5.5 at 47.4%, across 9 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Agents & tools
Index weight
Reference only
Models
9
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Given a vulnerable target, the agent must produce a working exploit, scored on capability.

How it is scored

Mean capability score, as published.

What to keep in mind

A dual-use capability measure; small model set.

9 of 9
#
1Claude Mythos PreviewAnthropic73.8%2026-04-07
2GPT-5.5OpenAI47.4%63.22026-04-23
3Claude Opus 4.7Anthropic26.5%64.52026-04-16
4Gemini 3.1 ProGoogle26.1%63.92026-02-19
5Claude Sonnet 4.6Anthropic23.6%59.02026-02-17
6Kimi K2.6Moonshot AIopen ↗18.4%60.52026-04-20
7GLM-5.1Zhipu AIopen ↗18.1%56.92026-04-07
8Claude Haiku 4.5Anthropic13.7%48.12025-10-15
9MiniMax-M2.7MiniMaxopen ↗13.3%56.12026-03-18

Cite as: BenchLeader, “ExploitBench leaderboard”, https://www.benchleader.com/benchmarks/exploitbench, data as of 19 Sept 2026.

ExploitBench: questions

What does ExploitBench measure?
Given a vulnerable target, the agent must produce a working exploit, scored on capability. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads ExploitBench?
Claude Mythos Preview leads ExploitBench with 73.8% as of 19 Sept 2026, ahead of GPT-5.5 at 47.4%.
How many models have ExploitBench results?
9 model configurations have a ExploitBench result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs ExploitBench and how often is it updated?
ExploitBench is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does ExploitBench count toward the BenchLeader Index?
No. ExploitBench is shown for reference but left out of the composite index.