BenchLeader

CUA-bench

Playing six commercial video games with only a keyboard and a mouse; half are held out. Run by Vals AI.

As of 19 Sept 2026, GPT-6 Astra leads CUA-bench on BenchLeader with 19.2%, ahead of Claude Fable 5.1 at 13.2%, across 5 model configurations with a published result.

Published by
Vals AI
Category
Agents & tools
Index weight
Reference only
Models
5
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

A computer-use agent is put in front of six commercial games with screen, keyboard and mouse and scored on progress. Three games are public, three held out.

How it is scored

Progress score across games, run by Vals AI.

What to keep in mind

Games test perception and control loops more than knowledge; scaffolding and frame rate matter.

5 of 5
#
1GPT-6 AstramaxOpenAI19.2%71.82026-09-18
2Claude Fable 5.1maxAnthropic13.2%69.72026-09-18
3Claude Opus 5maxAnthropic9.0%69.92026-09-18
4GPT-5.6 SolmaxOpenAI8.3%68.82026-09-18
5Gemini 3.8 FlashhighGoogle4.2%64.52026-09-18

Cite as: BenchLeader, “CUA-bench leaderboard”, https://www.benchleader.com/benchmarks/vals_cua_bench, data as of 19 Sept 2026.

CUA-bench: questions

What does CUA-bench measure?
A computer-use agent is put in front of six commercial games with screen, keyboard and mouse and scored on progress. Three games are public, three held out. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads CUA-bench?
GPT-6 Astra leads CUA-bench with 19.2% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 13.2%.
How many models have CUA-bench results?
5 model configurations have a CUA-bench result on BenchLeader, all taken from Vals AI.
Who runs CUA-bench and how often is it updated?
CUA-bench is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does CUA-bench count toward the BenchLeader Index?
No. CUA-bench is shown for reference but left out of the composite index.