BenchLeader

ProgramBench

Rebuilding programs from compiled binaries. Run by Vals AI.

As of 19 Sept 2026, Claude Fable 5.1 leads ProgramBench on BenchLeader with 7.0%, ahead of GPT-6 Astra at 5.5%, across 45 model configurations with a published result.

Published by
Vals AI
Category
Coding
Index weight
Reference only
Models
45
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

Given a binary, the model must recover working source that reproduces its behaviour, checked by tests.

How it is scored

Percent of programs rebuilt, run by Vals AI.

What to keep in mind

Reverse engineering is niche; scores are low for most models.

45 of 45
#
1Claude Fable 5.1Anthropic7.0%64.52026-09-16
2GPT-6 AstramaxOpenAI5.5%71.82026-09-16
3Claude Opus 5Anthropic3.0%66.72026-09-16
4Claude Fable 5Anthropic2.0%68.32026-09-16
5Kimi K3Moonshot AIopen ↗2.0%64.22026-09-16
6GPT-5.6 SolmaxOpenAI1.5%68.82026-09-16
7GLM-5.3maxZhipu AIopen ↗1.5%65.82026-09-16
8Gemini 3.8 FlashhighGoogle1.0%64.52026-09-16
9Claude Opus 4.8Anthropic1.0%61.92026-09-16
10GPT-5.5xhighOpenAI0.5%67.62026-09-16
11GPT-5.6 TerramaxOpenAI0.5%65.02026-09-16
12GLM-5.2maxZhipu AIopen ↗0.5%63.92026-09-16
13Claude Sonnet 4.6Anthropic0.5%59.02026-09-16
14GPT-5.4highOpenAI0.5%59.02026-09-16
15Inkling SmallThinking Machinesopen ↗0.5%56.32026-09-16
16Qwen3 8maxAlibaba0.0%66.52026-09-16
17GPT-5.4xhighOpenAI0.0%65.42026-09-16
18Gemini 3.7 FlashhighGoogle0.0%64.62026-09-16
19Claude Opus 4.7Anthropic0.0%64.52026-09-16
20DeepSeek V4 PromaxDeepSeekopen ↗0.0%64.22026-09-16
21Gemini 3.5 FlashhighGoogle0.0%63.62026-09-16
22Muse Spark 1.1xhighMeta0.0%61.82026-09-16
23Gemini 3.6 FlashhighGoogle0.0%61.72026-09-16
24Grok 4.5highxAI0.0%61.02026-09-16
25Gemini 3.1 ProhighGoogle0.0%60.72026-09-16
26Kimi K2.6Moonshot AIopen ↗0.0%60.52026-09-16
27GPT-5.6 LunamaxOpenAI0.0%60.22026-09-16
28Qwen3.8 27BxhighAlibabaopen ↗0.0%59.22026-09-16
29Gemini 3 FlashhighGoogle0.0%58.72026-09-16
30Qwen3.6 PlusAlibabaopen0.0%58.52026-09-16
31Grok 4.3highxAI0.0%58.22026-09-16
32Claude Sonnet 5Anthropic0.0%57.22026-09-16
33GLM-5.1Zhipu AIopen ↗0.0%56.92026-09-16
34MiniMax-M2.7MiniMaxopen ↗0.0%56.12026-09-16
35DeepSeek V4 Pro 0813maxDeepSeekopen ↗0.0%56.02026-09-16
36Kimi K2.7 CodeMoonshot AIopen ↗0.0%55.92026-09-16
37GPT-5.4 minixhighOpenAI0.0%55.72026-09-16
38GLM-5.3-FlashmaxZhipu AIopen ↗0.0%55.32026-09-16
39DeepSeek V4 Flash 0731highDeepSeekopen ↗0.0%55.02026-09-16
40Gemini 3.1 Flash LitehighGoogle0.0%49.12026-09-16
41Nemotron 3 Ultra 550B A55BNVIDIAopen ↗0.0%48.12026-09-16
42Claude Haiku 4.5thinkingAnthropic0.0%47.52026-09-16
43Gemini 3.5 Flash LitehighGoogle0.0%44.32026-09-16
44Laguna M.1Poolsideopen0.0%37.72026-09-16
45Laguna XS.2Poolside0.0%36.52026-09-16

Cite as: BenchLeader, “ProgramBench leaderboard”, https://www.benchleader.com/benchmarks/vals_programbench, data as of 19 Sept 2026.

ProgramBench: questions

What does ProgramBench measure?
Given a binary, the model must recover working source that reproduces its behaviour, checked by tests. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads ProgramBench?
Claude Fable 5.1 leads ProgramBench with 7.0% as of 19 Sept 2026, ahead of GPT-6 Astra at 5.5%.
How many models have ProgramBench results?
45 model configurations have a ProgramBench result on BenchLeader, all taken from Vals AI.
Who runs ProgramBench and how often is it updated?
ProgramBench is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does ProgramBench count toward the BenchLeader Index?
No. ProgramBench is shown for reference but left out of the composite index.