Vibe Code Bench 1-100
Extending a working web app, request by request, one hundred times. Run by Vals AI.
As of 19 Sept 2026, Claude Opus 5 leads Vibe Code Bench 1-100 on BenchLeader with 28.5%, ahead of Claude Fable 5.1 at 28.0%, across 18 model configurations with a published result.
- Published by
- Vals AI
- Category
- Coding
- Index weight
- Reference only
- Models
- 18
- Data as of
- 19 Sept 2026
Vals AI (vals.ai).
What the test looks like
Starting from a working app, the agent applies a hundred successive feature requests; the score is how many it completes without breaking the app.
How it is scored
Requests completed, run by Vals AI.
What to keep in mind
Tests sustained agentic coding rather than one-shot generation.
- 1Claude Opus 528.5%
- 2Claude Fable 5.128.0%
- 3GPT-6 Astra (max)27.6%
- 4GPT-5.6 Luna (max)22.6%
- 5Muse Spark 1.3 (max)20.5%
- 6GPT-5.6 Sol (max)20.0%
- 7GLM-5.3 (max)20.0%
- 8Gemini 3.8 Flash (high)18.8%
- 9Kimi K3 (max)18.2%
- 10DeepSeek V4 Pro 0813 (max)17.5%
- 11DeepSeek V4.1 Flash (high)16.4%
- 12GLM-5.3-Flash (max)16.0%
- 13GPT-5.6 Terra (max)14.8%
- 14Grok 4.6 (high)14.8%
- 15Claude Sonnet 513.8%
18 of 18
| # | ||||
|---|---|---|---|---|
| 1 | 28.5% | 66.7 | 2026-09-16 | |
| 2 | 28.0% | 64.5 | 2026-09-16 | |
| 3 | 27.6% | 71.8 | 2026-09-16 | |
| 4 | 22.6% | 60.2 | 2026-09-16 | |
| 5 | 20.5% | 69.3 | 2026-09-16 | |
| 6 | 20.0% | 68.8 | 2026-09-16 | |
| 7 | 20.0% | 65.8 | 2026-09-16 | |
| 8 | 18.8% | 64.5 | 2026-09-16 | |
| 9 | 18.2% | 67.2 | 2026-09-16 | |
| 10 | 17.5% | 56.0 | 2026-09-16 | |
| 11 | 16.4% | – | 2026-09-16 | |
| 12 | 16.0% | 55.3 | 2026-09-16 | |
| 13 | 14.8% | 65.0 | 2026-09-16 | |
| 14 | 14.8% | 64.5 | 2026-09-16 | |
| 15 | 13.8% | 57.2 | 2026-09-16 | |
| 16 | 12.8% | 66.5 | 2026-09-16 | |
| 17 | 9.2% | 57.5 | 2026-09-16 | |
| 18 | 6.7% | 60.7 | 2026-09-16 |
Cite as: BenchLeader, “Vibe Code Bench 1-100 leaderboard”, https://www.benchleader.com/benchmarks/vals_vibe_code_bench, data as of 19 Sept 2026.
Vibe Code Bench 1-100: questions
- What does Vibe Code Bench 1-100 measure?
- Starting from a working app, the agent applies a hundred successive feature requests; the score is how many it completes without breaking the app. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads Vibe Code Bench 1-100?
- Claude Opus 5 leads Vibe Code Bench 1-100 with 28.5% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 28.0%.
- How many models have Vibe Code Bench 1-100 results?
- 18 model configurations have a Vibe Code Bench 1-100 result on BenchLeader, all taken from Vals AI.
- Who runs Vibe Code Bench 1-100 and how often is it updated?
- Vibe Code Bench 1-100 is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does Vibe Code Bench 1-100 count toward the BenchLeader Index?
- No. Vibe Code Bench 1-100 is shown for reference but left out of the composite index.