ProgramBench
Rebuilding programs from compiled binaries. Run by Vals AI.
As of 19 Sept 2026, Claude Fable 5.1 leads ProgramBench on BenchLeader with 7.0%, ahead of GPT-6 Astra at 5.5%, across 45 model configurations with a published result.
- Published by
- Vals AI
- Category
- Coding
- Index weight
- Reference only
- Models
- 45
- Data as of
- 19 Sept 2026
Vals AI (vals.ai).
What the test looks like
Given a binary, the model must recover working source that reproduces its behaviour, checked by tests.
How it is scored
Percent of programs rebuilt, run by Vals AI.
What to keep in mind
Reverse engineering is niche; scores are low for most models.
- 1Claude Fable 5.17.0%
- 2GPT-6 Astra (max)5.5%
- 3Claude Opus 53.0%
- 4Claude Fable 52.0%
- 5Kimi K32.0%
- 6GPT-5.6 Sol (max)1.5%
- 7GLM-5.3 (max)1.5%
- 8Gemini 3.8 Flash (high)1.0%
- 9Claude Opus 4.81.0%
- 10GPT-5.5 (xhigh)0.5%
- 11GPT-5.6 Terra (max)0.5%
- 12GLM-5.2 (max)0.5%
- 13Claude Sonnet 4.60.5%
- 14GPT-5.4 (high)0.5%
- 15Inkling Small0.5%
45 of 45
| # | ||||
|---|---|---|---|---|
| 1 | 7.0% | 64.5 | 2026-09-16 | |
| 2 | 5.5% | 71.8 | 2026-09-16 | |
| 3 | 3.0% | 66.7 | 2026-09-16 | |
| 4 | 2.0% | 68.3 | 2026-09-16 | |
| 5 | 2.0% | 64.2 | 2026-09-16 | |
| 6 | 1.5% | 68.8 | 2026-09-16 | |
| 7 | 1.5% | 65.8 | 2026-09-16 | |
| 8 | 1.0% | 64.5 | 2026-09-16 | |
| 9 | 1.0% | 61.9 | 2026-09-16 | |
| 10 | 0.5% | 67.6 | 2026-09-16 | |
| 11 | 0.5% | 65.0 | 2026-09-16 | |
| 12 | 0.5% | 63.9 | 2026-09-16 | |
| 13 | 0.5% | 59.0 | 2026-09-16 | |
| 14 | 0.5% | 59.0 | 2026-09-16 | |
| 15 | Inkling SmallThinking Machinesopen ↗ | 0.5% | 56.3 | 2026-09-16 |
| 16 | 0.0% | 66.5 | 2026-09-16 | |
| 17 | 0.0% | 65.4 | 2026-09-16 | |
| 18 | 0.0% | 64.6 | 2026-09-16 | |
| 19 | 0.0% | 64.5 | 2026-09-16 | |
| 20 | 0.0% | 64.2 | 2026-09-16 | |
| 21 | 0.0% | 63.6 | 2026-09-16 | |
| 22 | 0.0% | 61.8 | 2026-09-16 | |
| 23 | 0.0% | 61.7 | 2026-09-16 | |
| 24 | 0.0% | 61.0 | 2026-09-16 | |
| 25 | 0.0% | 60.7 | 2026-09-16 | |
| 26 | 0.0% | 60.5 | 2026-09-16 | |
| 27 | 0.0% | 60.2 | 2026-09-16 | |
| 28 | 0.0% | 59.2 | 2026-09-16 | |
| 29 | 0.0% | 58.7 | 2026-09-16 | |
| 30 | 0.0% | 58.5 | 2026-09-16 | |
| 31 | 0.0% | 58.2 | 2026-09-16 | |
| 32 | 0.0% | 57.2 | 2026-09-16 | |
| 33 | 0.0% | 56.9 | 2026-09-16 | |
| 34 | 0.0% | 56.1 | 2026-09-16 | |
| 35 | 0.0% | 56.0 | 2026-09-16 | |
| 36 | 0.0% | 55.9 | 2026-09-16 | |
| 37 | 0.0% | 55.7 | 2026-09-16 | |
| 38 | 0.0% | 55.3 | 2026-09-16 | |
| 39 | 0.0% | 55.0 | 2026-09-16 | |
| 40 | 0.0% | 49.1 | 2026-09-16 | |
| 41 | 0.0% | 48.1 | 2026-09-16 | |
| 42 | 0.0% | 47.5 | 2026-09-16 | |
| 43 | 0.0% | 44.3 | 2026-09-16 | |
| 44 | 0.0% | 37.7 | 2026-09-16 | |
| 45 | 0.0% | 36.5 | 2026-09-16 |
Cite as: BenchLeader, “ProgramBench leaderboard”, https://www.benchleader.com/benchmarks/vals_programbench, data as of 19 Sept 2026.
ProgramBench: questions
- What does ProgramBench measure?
- Given a binary, the model must recover working source that reproduces its behaviour, checked by tests. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads ProgramBench?
- Claude Fable 5.1 leads ProgramBench with 7.0% as of 19 Sept 2026, ahead of GPT-6 Astra at 5.5%.
- How many models have ProgramBench results?
- 45 model configurations have a ProgramBench result on BenchLeader, all taken from Vals AI.
- Who runs ProgramBench and how often is it updated?
- ProgramBench is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does ProgramBench count toward the BenchLeader Index?
- No. ProgramBench is shown for reference but left out of the composite index.