VTB
Scale's vision task benchmark across practical image tasks. Scale AI.
As of 19 Sept 2026, Muse Spark 1.1 leads VTB on BenchLeader with 44.8%, ahead of GPT-5.4 at 29.2%, across 21 model configurations with a published result.
- Published by
- Scale AI SEAL
- Category
- Multimodal
- Index weight
- Reference only
- Models
- 21
- Data as of
- 19 Sept 2026
Scale AI SEAL Leaderboards.
What the test looks like
Practical vision tasks, from reading charts to spotting differences, graded against references.
How it is scored
Accuracy, published by Scale AI.
What to keep in mind
Vision models only.
- 1Muse Spark 1.144.8%
- 2GPT-5.4 (high)29.2%
- 3Gemini 3.1 Pro29.0%
- 4Claude Opus 4.6 (thinking)27.5%
- 5Gemini 3 Pro26.9%
- 6GPT-5 (thinking)18.7%
- 7GPT-517.0%
- 8o313.7%
- 9Gemini 2.5 Pro11.8%
- 10o4-mini11.1%
- 11Claude Sonnet 4.5 (thinking)6.2%
- 12Claude Sonnet 4.55.6%
- 13GPT-4.15.5%
- 14Claude Opus 4.1 (thinking)5.2%
- 15Claude Opus 4.14.7%
21 of 21
| # | ||||
|---|---|---|---|---|
| 1 | 44.8% | 65.2 | 2026-07-09 | |
| 2 | 29.2% | 59.0 | 2026-03-24 | |
| 3 | 29.0% | 63.9 | 2026-03-24 | |
| 4 | 27.5% | 62.5 | 2026-03-24 | |
| 5 | 26.9% | 61.1 | 2025-11-20 | |
| 6 | 18.7% | – | 2025-10-08 | |
| 7 | 17.0% | 57.2 | 2025-10-08 | |
| 8 | 13.7% | 60.3 | 2025-10-08 | |
| 9 | 11.8% | 54.5 | 2025-10-08 | |
| 10 | 11.1% | 57.4 | 2025-10-08 | |
| 11 | 6.2% | 53.4 | 2025-10-08 | |
| 12 | 5.6% | 54.3 | 2025-10-08 | |
| 13 | 5.5% | 49.0 | 2025-10-08 | |
| 14 | 5.2% | 57.2 | 2025-10-08 | |
| 15 | 4.7% | 51.9 | 2025-10-08 | |
| 16 | 4.7% | 52.5 | 2025-10-08 | |
| 17 | 4.5% | 49.7 | 2025-10-08 | |
| 18 | 4.4% | 54.0 | 2025-10-08 | |
| 19 | 2.0% | 45.7 | 2025-10-08 | |
| 20 | 1.6% | 40.2 | 2025-10-08 | |
| 21 | 1.4% | 43.4 | 2025-10-08 |
Cite as: BenchLeader, “VTB leaderboard”, https://www.benchleader.com/benchmarks/scale_vtb, data as of 19 Sept 2026.
VTB: questions
- What does VTB measure?
- Practical vision tasks, from reading charts to spotting differences, graded against references. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads VTB?
- Muse Spark 1.1 leads VTB with 44.8% as of 19 Sept 2026, ahead of GPT-5.4 at 29.2%.
- How many models have VTB results?
- 21 model configurations have a VTB result on BenchLeader, all taken from Scale AI SEAL.
- Who runs VTB and how often is it updated?
- VTB is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does VTB count toward the BenchLeader Index?
- No. VTB is shown for reference but left out of the composite index.