Vibe Code Bench v1.1
Building web apps from scratch from a brief. Run by Vals AI.
As of 19 Sept 2026, Claude Fable 5 leads Vibe Code Bench v1.1 on BenchLeader with 90.3%, ahead of Claude Fable 5.1 at 90.3%, across 93 model configurations with a published result.
- Published by
- Vals AI
- Category
- Coding
- Index weight
- Reference only
- Models
- 93
- Data as of
- 19 Sept 2026
Vals AI (vals.ai).
What the test looks like
From a product brief the agent builds a working web app, judged on function and completeness.
How it is scored
Rubric score, run by Vals AI.
What to keep in mind
Rubric-graded; design choices and framework familiarity affect scores.
- 1Claude Fable 590.3%
- 2Claude Fable 5.190.3%
- 3GPT-6 Astra (max)89.6%
- 4Claude Opus 588.4%
- 5Muse Spark 1.3 (max)85.9%
- 6Kimi K385.0%
- 7DeepSeek V4.1 Flash (high)84.7%
- 8Muse Spark 1.3 (xhigh)82.9%
- 9Claude Opus 4.882.7%
- 10DeepSeek V4 Pro 0813 (max)82.3%
- 11Claude Sonnet 581.3%
- 12GPT-5.6 Sol (max)80.5%
- 13Muse Spark 1.2 (xhigh)79.1%
- 14Gemini 3.8 Flash (high)78.7%
- 15GLM-5.3 (max)78.1%
93 of 93
| # | ||||
|---|---|---|---|---|
| 1 | 90.3% | 68.3 | 2026-09-15 | |
| 2 | 90.3% | 64.5 | 2026-09-15 | |
| 3 | 89.6% | 71.8 | 2026-09-15 | |
| 4 | 88.4% | 66.7 | 2026-09-15 | |
| 5 | 85.9% | 69.3 | 2026-09-15 | |
| 6 | 85.0% | 64.2 | 2026-09-15 | |
| 7 | 84.7% | – | 2026-09-15 | |
| 8 | 82.9% | 68.0 | 2026-09-15 | |
| 9 | 82.7% | 61.9 | 2026-09-15 | |
| 10 | 82.3% | 56.0 | 2026-09-15 | |
| 11 | 81.3% | 57.2 | 2026-09-15 | |
| 12 | 80.5% | 68.8 | 2026-09-15 | |
| 13 | 79.1% | 64.1 | 2026-09-15 | |
| 14 | 78.7% | 64.5 | 2026-09-15 | |
| 15 | 78.1% | 65.8 | 2026-09-15 | |
| 16 | 77.5% | – | 2026-09-15 | |
| 17 | 77.1% | 60.2 | 2026-09-15 | |
| 18 | 76.2% | 64.5 | 2026-09-15 | |
| 19 | 74.7% | 55.0 | 2026-09-15 | |
| 20 | 74.6% | 65.0 | 2026-09-15 | |
| 21 | 72.2% | 61.8 | 2026-09-15 | |
| 22 | 71.0% | 64.5 | 2026-09-15 | |
| 23 | 70.4% | 64.6 | 2026-09-15 | |
| 24 | 69.8% | 67.6 | 2026-09-15 | |
| 25 | 69.0% | 61.0 | 2026-09-15 | |
| 26 | 67.4% | 65.4 | 2026-09-15 | |
| 27 | 67.4% | – | 2026-09-15 | |
| 28 | 64.8% | 59.2 | 2026-09-15 | |
| 29 | 64.7% | 66.5 | 2026-09-15 | |
| 30 | 64.0% | 61.7 | 2026-09-15 | |
| 31 | 64.0% | 63.9 | 2026-09-15 | |
| 32 | 61.8% | 65.4 | 2026-09-15 | |
| 33 | 58.2% | – | 2026-09-15 | |
| 34 | 57.6% | 63.7 | 2026-09-15 | |
| 35 | 55.8% | – | 2026-09-15 | |
| 36 | 53.5% | 62.5 | 2026-09-15 | |
| 37 | 53.5% | 62.0 | 2026-09-15 | |
| 38 | 51.5% | 59.0 | 2026-09-15 | |
| 39 | 49.9% | 64.2 | 2026-09-15 | |
| 40 | 49.6% | – | 2026-09-15 | |
| 41 | 48.7% | 63.6 | 2026-09-15 | |
| 42 | 48.5% | 59.2 | 2026-09-15 | |
| 43 | 48.0% | 55.7 | 2026-09-15 | |
| 44 | 47.7% | 61.0 | 2026-09-15 | |
| 45 | 47.6% | 57.5 | 2026-09-15 | |
| 46 | 47.2% | 55.9 | 2026-09-15 | |
| 47 | 46.4% | 59.9 | 2026-09-15 | |
| 48 | 37.9% | – | 2026-09-15 | |
| 49 | 37.9% | 60.5 | 2026-09-15 | |
| 50 | 37.2% | 44.3 | 2026-09-15 | |
| 51 | 34.1% | 59.2 | 2026-09-15 | |
| 52 | 32.0% | 60.7 | 2026-09-15 | |
| 53 | 31.5% | 56.9 | 2026-09-15 | |
| 54 | 30.8% | 55.3 | 2026-09-15 | |
| 55 | 26.1% | 49.1 | 2026-09-15 | |
| 56 | 25.6% | 58.5 | 2026-09-15 | |
| 57 | 24.6% | 58.7 | 2026-09-15 | |
| 58 | 23.4% | 59.6 | 2026-09-15 | |
| 59 | 22.6% | 53.4 | 2026-09-15 | |
| 60 | 22.2% | – | 2026-09-15 |
Cite as: BenchLeader, “Vibe Code Bench v1.1 leaderboard”, https://www.benchleader.com/benchmarks/vals_vibe_code, data as of 19 Sept 2026.
Vibe Code Bench v1.1: questions
- What does Vibe Code Bench v1.1 measure?
- From a product brief the agent builds a working web app, judged on function and completeness. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads Vibe Code Bench v1.1?
- Claude Fable 5 leads Vibe Code Bench v1.1 with 90.3% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 90.3%.
- How many models have Vibe Code Bench v1.1 results?
- 93 model configurations have a Vibe Code Bench v1.1 result on BenchLeader, all taken from Vals AI.
- Who runs Vibe Code Bench v1.1 and how often is it updated?
- Vibe Code Bench v1.1 is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does Vibe Code Bench v1.1 count toward the BenchLeader Index?
- No. Vibe Code Bench v1.1 is shown for reference but left out of the composite index.