Blueprint-Bench 2
Turning photos of rooms into floor plans.
As of 19 Sept 2026, Claude Fable 5.1 leads Blueprint-Bench 2 on BenchLeader with 41.9%, ahead of Claude Fable 5 at 38.6%, across 24 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Multimodal
- Index weight
- Reference only
- Models
- 24
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
From photographs of a home the model must draw the floor plan, compared with the real one.
How it is scored
Score, as published.
What to keep in mind
Vision models only; small set.
- 1Claude Fable 5.141.9%
- 2Claude Fable 538.6%
- 3GPT-5.536.2%
- 4GPT-5.6 Sol33.6%
- 5Gemini 3.5 Flash33.6%
- 6Grok 4.633.2%
- 7Gemini 3.6 Flash31.2%
- 8GPT-5.6 Terra30.8%
- 9Claude Opus 530.4%
- 10Kimi K329.5%
- 11Grok 4.527.3%
- 12GPT-5.427.1%
- 13Gemini 3.1 Pro26.5%
- 14Claude Sonnet 524.9%
- 15Claude Opus 4.724.5%
24 of 24
| # | ||||
|---|---|---|---|---|
| 1 | 41.9% | 64.5 | 2026-09-01 | |
| 2 | 38.6% | 68.3 | 2026-06-09 | |
| 3 | 36.2% | 63.2 | 2026-04-23 | |
| 4 | 33.6% | 56.8 | 2026-07-09 | |
| 5 | 33.6% | 51.3 | 2026-05-19 | |
| 6 | 33.2% | 59.3 | 2026-08-12 | |
| 7 | 31.2% | 51.4 | 2026-07-21 | |
| 8 | 30.8% | – | 2026-07-09 | |
| 9 | 30.4% | 66.7 | 2026-07-24 | |
| 10 | 29.5% | 64.2 | 2026-07-16 | |
| 11 | 27.3% | 60.3 | 2026-07-08 | |
| 12 | 27.1% | 59.2 | 2026-03-05 | |
| 13 | 26.5% | 63.9 | 2026-02-19 | |
| 14 | 24.9% | 57.2 | 2026-06-30 | |
| 15 | 24.5% | 64.5 | 2026-04-16 | |
| 16 | 22.6% | – | 2026-07-09 | |
| 17 | 14.5% | 61.9 | 2026-05-28 | |
| 18 | 6.7% | 59.0 | 2026-02-17 | |
| 19 | 3.9% | 60.5 | 2026-04-20 | |
| 20 | 0.0% | 57.7 | 2025-12-17 | |
| 21 | 0.0% | 50.5 | 2026-04-17 | |
| 22 | 0.0% | 48.1 | 2025-10-15 | |
| 23 | 0.0% | – | 2026-02-17 | |
| 24 | 0.0% | – | 2026-04-14 |
Cite as: BenchLeader, “Blueprint-Bench 2 leaderboard”, https://www.benchleader.com/benchmarks/blueprint_bench_2, data as of 19 Sept 2026.
Blueprint-Bench 2: questions
- What does Blueprint-Bench 2 measure?
- From photographs of a home the model must draw the floor plan, compared with the real one. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads Blueprint-Bench 2?
- Claude Fable 5.1 leads Blueprint-Bench 2 with 41.9% as of 19 Sept 2026, ahead of Claude Fable 5 at 38.6%.
- How many models have Blueprint-Bench 2 results?
- 24 model configurations have a Blueprint-Bench 2 result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs Blueprint-Bench 2 and how often is it updated?
- Blueprint-Bench 2 is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does Blueprint-Bench 2 count toward the BenchLeader Index?
- No. Blueprint-Bench 2 is shown for reference but left out of the composite index.