BenchLeader

Blueprint-Bench 2

Turning photos of rooms into floor plans.

As of 19 Sept 2026, Claude Fable 5.1 leads Blueprint-Bench 2 on BenchLeader with 41.9%, ahead of Claude Fable 5 at 38.6%, across 24 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Multimodal
Index weight
Reference only
Models
24
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

From photographs of a home the model must draw the floor plan, compared with the real one.

How it is scored

Score, as published.

What to keep in mind

Vision models only; small set.

24 of 24
#
1Claude Fable 5.1Anthropic41.9%64.52026-09-01
2Claude Fable 5Anthropic38.6%68.32026-06-09
3GPT-5.5OpenAI36.2%63.22026-04-23
4GPT-5.6 SolOpenAI33.6%56.82026-07-09
5Gemini 3.5 FlashGoogle33.6%51.32026-05-19
6Grok 4.6xAI33.2%59.32026-08-12
7Gemini 3.6 FlashGoogle31.2%51.42026-07-21
8GPT-5.6 TerraOpenAI30.8%2026-07-09
9Claude Opus 5Anthropic30.4%66.72026-07-24
10Kimi K3Moonshot AIopen ↗29.5%64.22026-07-16
11Grok 4.5xAI27.3%60.32026-07-08
12GPT-5.4OpenAI27.1%59.22026-03-05
13Gemini 3.1 ProGoogle26.5%63.92026-02-19
14Claude Sonnet 5Anthropic24.9%57.22026-06-30
15Claude Opus 4.7Anthropic24.5%64.52026-04-16
16GPT-5.6 LunaOpenAI22.6%2026-07-09
17Claude Opus 4.8Anthropic14.5%61.92026-05-28
18Claude Sonnet 4.6Anthropic6.7%59.02026-02-17
19Kimi K2.6Moonshot AIopen ↗3.9%60.52026-04-20
20Gemini 3 FlashGoogle0.0%57.72025-12-17
21Grok 4.3xAI0.0%50.52026-04-17
22Claude Haiku 4.5Anthropic0.0%48.12025-10-15
23Grok 4.20xAI0.0%2026-02-17
24Gemini Robotics-ER 1.6Google0.0%2026-04-14

Cite as: BenchLeader, “Blueprint-Bench 2 leaderboard”, https://www.benchleader.com/benchmarks/blueprint_bench_2, data as of 19 Sept 2026.

Blueprint-Bench 2: questions

What does Blueprint-Bench 2 measure?
From photographs of a home the model must draw the floor plan, compared with the real one. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads Blueprint-Bench 2?
Claude Fable 5.1 leads Blueprint-Bench 2 with 41.9% as of 19 Sept 2026, ahead of Claude Fable 5 at 38.6%.
How many models have Blueprint-Bench 2 results?
24 model configurations have a Blueprint-Bench 2 result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs Blueprint-Bench 2 and how often is it updated?
Blueprint-Bench 2 is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does Blueprint-Bench 2 count toward the BenchLeader Index?
No. Blueprint-Bench 2 is shown for reference but left out of the composite index.