BenchLeader

Vibe Code Bench v1.1

Building web apps from scratch from a brief. Run by Vals AI.

As of 19 Sept 2026, Claude Fable 5 leads Vibe Code Bench v1.1 on BenchLeader with 90.3%, ahead of Claude Fable 5.1 at 90.3%, across 93 model configurations with a published result.

Published by
Vals AI
Category
Coding
Index weight
Reference only
Models
93
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

From a product brief the agent builds a working web app, judged on function and completeness.

How it is scored

Rubric score, run by Vals AI.

What to keep in mind

Rubric-graded; design choices and framework familiarity affect scores.

93 of 93
#
1Claude Fable 5Anthropic90.3%68.32026-09-15
2Claude Fable 5.1Anthropic90.3%64.52026-09-15
3GPT-6 AstramaxOpenAI89.6%71.82026-09-15
4Claude Opus 5Anthropic88.4%66.72026-09-15
5Muse Spark 1.3maxMeta85.9%69.32026-09-15
6Kimi K3Moonshot AIopen ↗85.0%64.22026-09-15
7DeepSeek V4.1 FlashhighDeepSeekopen ↗84.7%2026-09-15
8Muse Spark 1.3xhighMeta82.9%68.02026-09-15
9Claude Opus 4.8Anthropic82.7%61.92026-09-15
10DeepSeek V4 Pro 0813maxDeepSeekopen ↗82.3%56.02026-09-15
11Claude Sonnet 5Anthropic81.3%57.22026-09-15
12GPT-5.6 SolmaxOpenAI80.5%68.82026-09-15
13Muse Spark 1.2xhighMeta79.1%64.12026-09-15
14Gemini 3.8 FlashhighGoogle78.7%64.52026-09-15
15GLM-5.3maxZhipu AIopen ↗78.1%65.82026-09-15
16Claude Opus 4.8 Claude CodeAnthropic77.5%2026-09-15
17GPT-5.6 LunamaxOpenAI77.1%60.22026-09-15
18Grok 4.6highxAI76.2%64.52026-09-15
19DeepSeek V4 Flash 0731highDeepSeekopen ↗74.7%55.02026-09-15
20GPT-5.6 TerramaxOpenAI74.6%65.02026-09-15
21Muse Spark 1.1xhighMeta72.2%61.82026-09-15
22Claude Opus 4.7Anthropic71.0%64.52026-09-15
23Gemini 3.7 FlashhighGoogle70.4%64.62026-09-15
24GPT-5.5xhighOpenAI69.8%67.62026-09-15
25Grok 4.5highxAI69.0%61.02026-09-15
26GPT-5.4xhighOpenAI67.4%65.42026-09-15
27GPT 5.5 FactoryOpenAI67.4%2026-09-15
28Qwen3.8 27BxhighAlibabaopen ↗64.8%59.22026-09-15
29Qwen3 8maxAlibaba64.7%66.52026-09-15
30Gemini 3.6 FlashhighGoogle64.0%61.72026-09-15
31GLM-5.2maxZhipu AIopen ↗64.0%63.92026-09-15
32GPT-5.3 CodexxhighOpenAI61.8%65.42026-09-15
33GPT 5.5 CodexOpenAI58.2%2026-09-15
34Claude Opus 4.6Anthropic57.6%63.72026-09-15
35Claude Sonnet 4.6 Claude CodeAnthropic55.8%2026-09-15
36Claude Opus 4.6thinkingAnthropic53.5%62.52026-09-15
37GPT-5.2xhighOpenAI53.5%62.02026-09-15
38Claude Sonnet 4.6Anthropic51.5%59.02026-09-15
39DeepSeek V4 PromaxDeepSeekopen ↗49.9%64.22026-09-15
40Composer 2.5Cursor49.6%2026-09-15
41Gemini 3.5 FlashhighGoogle48.7%63.62026-09-15
42GPT-5.4OpenAI48.5%59.22026-09-15
43GPT-5.4 minixhighOpenAI48.0%55.72026-09-15
44Qwen3 7maxAlibaba47.7%61.02026-09-15
45MiniMax-M3MiniMaxopen ↗47.6%57.52026-09-15
46Kimi K2.7 CodeMoonshot AIopen ↗47.2%55.92026-09-15
47Qwen3.7 PlusAlibaba46.4%59.92026-09-15
48GPT-5.2-CodexhighOpenAI37.9%2026-09-15
49Kimi K2.6Moonshot AIopen ↗37.9%60.52026-09-15
50Gemini 3.5 Flash LitehighGoogle37.2%44.32026-09-15
51MiMo-V2.5-ProXiaomiopen ↗34.1%59.22026-09-15
52Gemini 3.1 ProhighGoogle32.0%60.72026-09-15
53GLM-5.1Zhipu AIopen ↗31.5%56.92026-09-15
54GLM-5.3-FlashmaxZhipu AIopen ↗30.8%55.32026-09-15
55GPT-5.4 nanohighOpenAI26.1%49.12026-09-15
56Qwen3.6 PlusAlibabaopen25.6%58.52026-09-15
57GPT-5.1highOpenAI24.6%58.72026-09-15
58GLM-5thinkingZhipu AIopen23.4%59.62026-09-15
59Claude Sonnet 4.5thinkingAnthropic22.6%53.42026-09-15
60GPT-5.1-Codex-MaxmaxOpenAI22.2%2026-09-15

Cite as: BenchLeader, “Vibe Code Bench v1.1 leaderboard”, https://www.benchleader.com/benchmarks/vals_vibe_code, data as of 19 Sept 2026.

Vibe Code Bench v1.1: questions

What does Vibe Code Bench v1.1 measure?
From a product brief the agent builds a working web app, judged on function and completeness. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads Vibe Code Bench v1.1?
Claude Fable 5 leads Vibe Code Bench v1.1 with 90.3% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 90.3%.
How many models have Vibe Code Bench v1.1 results?
93 model configurations have a Vibe Code Bench v1.1 result on BenchLeader, all taken from Vals AI.
Who runs Vibe Code Bench v1.1 and how often is it updated?
Vibe Code Bench v1.1 is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does Vibe Code Bench v1.1 count toward the BenchLeader Index?
No. Vibe Code Bench v1.1 is shown for reference but left out of the composite index.