BenchLeader

CyberBench

Finding and patching security bugs in real code. Run by Vals AI.

As of 19 Sept 2026, GPT-5.6 Sol leads CyberBench on BenchLeader with 88.1%, ahead of GPT-5.6 Luna at 83.9%, across 24 model configurations with a published result.

Published by
Vals AI
Category
Agents & tools
Index weight
Reference only
Models
24
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

Given a codebase with a vulnerability, the agent must find it and patch it so the exploit no longer works and the tests still pass.

How it is scored

Percent of vulnerabilities fixed, run by Vals AI.

What to keep in mind

Complements Cybench's capture-the-flag framing with defensive work.

24 of 24
#
1GPT-5.6 SolOpenAI88.1%56.82026-09-16
2GPT-5.6 LunaOpenAI83.9%2026-09-16
3GPT-5.5OpenAI80.5%63.22026-09-16
4Kimi K3Moonshot AIopen ↗79.0%64.22026-09-16
5DeepSeek V4.1 FlashhighDeepSeekopen ↗78.8%2026-09-16
6DeepSeek V4 Flash 0731highDeepSeekopen ↗77.5%55.02026-09-16
7GLM-5.2Zhipu AIopen ↗77.4%52.12026-09-16
8MiniMax-M3MiniMaxopen ↗76.3%57.52026-09-16
9Kimi K2.6Moonshot AIopen ↗74.3%60.52026-09-16
10GPT-5.4OpenAI72.9%59.22026-09-16
11Gemini 3.5 Flash LitehighGoogle72.3%44.32026-09-16
12GLM-5.3-FlashmaxZhipu AIopen ↗71.5%55.32026-09-16
13Gemini 3.5 FlashGoogle69.3%51.32026-09-16
14Inkling SmallThinking Machinesopen ↗68.8%56.32026-09-16
15DeepSeek V4 ProDeepSeekopen ↗68.6%52.62026-09-16
16DeepSeek V4 FlashDeepSeekopen ↗66.7%51.62026-09-16
17Claude Opus 4.7Anthropic52.1%64.52026-09-16
18Claude Opus 4.8Anthropic50.9%61.92026-09-16
19Gemini 3.6 FlashhighGoogle48.5%61.72026-09-16
20Grok 4.3xAI47.5%50.52026-09-16
21Gemini 3.8 FlashhighGoogle43.8%64.52026-09-16
22Claude Opus 5Anthropic40.7%66.72026-09-16
23Qwen3.7 PlusAlibaba37.5%59.92026-09-16
24Gemini 3.1 ProGoogle36.4%63.92026-09-16

Cite as: BenchLeader, “CyberBench leaderboard”, https://www.benchleader.com/benchmarks/vals_cyber, data as of 19 Sept 2026.

CyberBench: questions

What does CyberBench measure?
Given a codebase with a vulnerability, the agent must find it and patch it so the exploit no longer works and the tests still pass. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads CyberBench?
GPT-5.6 Sol leads CyberBench with 88.1% as of 19 Sept 2026, ahead of GPT-5.6 Luna at 83.9%.
How many models have CyberBench results?
24 model configurations have a CyberBench result on BenchLeader, all taken from Vals AI.
Who runs CyberBench and how often is it updated?
CyberBench is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does CyberBench count toward the BenchLeader Index?
No. CyberBench is shown for reference but left out of the composite index.