BenchLeader

MindCube

Spatial reasoning from multiple views.

As of 19 Sept 2026, Gemma 3 12B leads MindCube on BenchLeader with 46.7%, ahead of mPLUG-Owl3-7B-241101 at 44.9%, across 5 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Multimodal
Index weight
Reference only
Models
5
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Questions that require building a mental map from several views of a scene.

How it is scored

Overall score, as published.

What to keep in mind

Very few models measured.

5 of 5
#
1Gemma 3 12BGoogleopen ↗46.7%37.72025-03-12
2mPLUG-Owl3-7B-241101Unknown44.9%2024-11-26
3Claude Sonnet 4Anthropic44.8%49.72025-05-22
4LLaVA-Video-7B-Qwen2OpenBMB42.0%2024-09-02
5LongVA-7BUnknown29.5%2024-06-13

Cite as: BenchLeader, “MindCube leaderboard”, https://www.benchleader.com/benchmarks/mindcube, data as of 19 Sept 2026.

MindCube: questions

What does MindCube measure?
Questions that require building a mental map from several views of a scene. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads MindCube?
Gemma 3 12B leads MindCube with 46.7% as of 19 Sept 2026, ahead of mPLUG-Owl3-7B-241101 at 44.9%.
How many models have MindCube results?
5 model configurations have a MindCube result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs MindCube and how often is it updated?
MindCube is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does MindCube count toward the BenchLeader Index?
No. MindCube is shown for reference but left out of the composite index.