CyberBench
Finding and patching security bugs in real code. Run by Vals AI.
As of 19 Sept 2026, GPT-5.6 Sol leads CyberBench on BenchLeader with 88.1%, ahead of GPT-5.6 Luna at 83.9%, across 24 model configurations with a published result.
- Published by
- Vals AI
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 24
- Data as of
- 19 Sept 2026
Vals AI (vals.ai).
What the test looks like
Given a codebase with a vulnerability, the agent must find it and patch it so the exploit no longer works and the tests still pass.
How it is scored
Percent of vulnerabilities fixed, run by Vals AI.
What to keep in mind
Complements Cybench's capture-the-flag framing with defensive work.
- 1GPT-5.6 Sol88.1%
- 2GPT-5.6 Luna83.9%
- 3GPT-5.580.5%
- 4Kimi K379.0%
- 5DeepSeek V4.1 Flash (high)78.8%
- 6DeepSeek V4 Flash 0731 (high)77.5%
- 7GLM-5.277.4%
- 8MiniMax-M376.3%
- 9Kimi K2.674.3%
- 10GPT-5.472.9%
- 11Gemini 3.5 Flash Lite (high)72.3%
- 12GLM-5.3-Flash (max)71.5%
- 13Gemini 3.5 Flash69.3%
- 14Inkling Small68.8%
- 15DeepSeek V4 Pro68.6%
24 of 24
| # | ||||
|---|---|---|---|---|
| 1 | 88.1% | 56.8 | 2026-09-16 | |
| 2 | 83.9% | – | 2026-09-16 | |
| 3 | 80.5% | 63.2 | 2026-09-16 | |
| 4 | 79.0% | 64.2 | 2026-09-16 | |
| 5 | 78.8% | – | 2026-09-16 | |
| 6 | 77.5% | 55.0 | 2026-09-16 | |
| 7 | 77.4% | 52.1 | 2026-09-16 | |
| 8 | 76.3% | 57.5 | 2026-09-16 | |
| 9 | 74.3% | 60.5 | 2026-09-16 | |
| 10 | 72.9% | 59.2 | 2026-09-16 | |
| 11 | 72.3% | 44.3 | 2026-09-16 | |
| 12 | 71.5% | 55.3 | 2026-09-16 | |
| 13 | 69.3% | 51.3 | 2026-09-16 | |
| 14 | Inkling SmallThinking Machinesopen ↗ | 68.8% | 56.3 | 2026-09-16 |
| 15 | 68.6% | 52.6 | 2026-09-16 | |
| 16 | 66.7% | 51.6 | 2026-09-16 | |
| 17 | 52.1% | 64.5 | 2026-09-16 | |
| 18 | 50.9% | 61.9 | 2026-09-16 | |
| 19 | 48.5% | 61.7 | 2026-09-16 | |
| 20 | 47.5% | 50.5 | 2026-09-16 | |
| 21 | 43.8% | 64.5 | 2026-09-16 | |
| 22 | 40.7% | 66.7 | 2026-09-16 | |
| 23 | 37.5% | 59.9 | 2026-09-16 | |
| 24 | 36.4% | 63.9 | 2026-09-16 |
Cite as: BenchLeader, “CyberBench leaderboard”, https://www.benchleader.com/benchmarks/vals_cyber, data as of 19 Sept 2026.
CyberBench: questions
- What does CyberBench measure?
- Given a codebase with a vulnerability, the agent must find it and patch it so the exploit no longer works and the tests still pass. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads CyberBench?
- GPT-5.6 Sol leads CyberBench with 88.1% as of 19 Sept 2026, ahead of GPT-5.6 Luna at 83.9%.
- How many models have CyberBench results?
- 24 model configurations have a CyberBench result on BenchLeader, all taken from Vals AI.
- Who runs CyberBench and how often is it updated?
- CyberBench is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does CyberBench count toward the BenchLeader Index?
- No. CyberBench is shown for reference but left out of the composite index.