BenchLeader

Terminal-Bench 4.0 (Vals)

Terminal-Bench 4.0, the frontier-difficulty set, in Vals AI's harness. Run by Vals AI.

As of 19 Sept 2026, GPT-6 Astra leads Terminal-Bench 4.0 (Vals) on BenchLeader with 57.1%, ahead of Claude Fable 5.1 at 49.5%, across 27 model configurations with a published result.

Published by
Vals AI
Category
Agents & tools
Index weight
Reference only
Models
27
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

Terminal-Bench 4.0 tasks written to defeat current frontier agents, run by Vals AI.

How it is scored

Percent of tasks passing, run by Vals AI.

What to keep in mind

New and hard; small set, low scores.

27 of 27
#
1GPT-6 AstramaxOpenAI57.1%71.82026-09-16
2Claude Fable 5.1Anthropic49.5%64.52026-09-16
3Claude Opus 5Anthropic45.5%66.72026-09-16
4Muse Spark 1.3maxMeta27.8%69.32026-09-16
5GPT-5.6 SolmaxOpenAI27.8%68.82026-09-16
6GPT-5.6 TerramaxOpenAI26.3%65.02026-09-16
7GLM-5.3maxZhipu AIopen ↗25.3%65.82026-09-16
8Qwen3 8maxAlibaba24.8%66.52026-09-16
9Claude Fable 5Anthropic22.7%68.32026-09-16
10GLM-5.3-FlashmaxZhipu AIopen ↗19.7%55.32026-09-16
11Grok 4.6highxAI17.2%64.52026-09-16
12Claude Opus 4.8Anthropic16.2%61.92026-09-16
13Muse Spark 1.3xhighMeta15.2%68.02026-09-16
14Gemini 3.8 FlashhighGoogle13.1%64.52026-09-16
15Kimi K3maxMoonshot AIopen ↗12.6%67.22026-09-16
16DeepSeek V4.1 FlashhighDeepSeekopen ↗11.6%2026-09-16
17DeepSeek V4 Flash 0731highDeepSeekopen ↗9.1%55.02026-09-16
18Claude Sonnet 5Anthropic8.1%57.22026-09-16
19Grok 4.5highxAI6.6%61.02026-09-16
20Gemini 3.7 FlashhighGoogle6.1%64.62026-09-16
21Muse Spark 1.2xhighMeta5.6%64.12026-09-16
22Gemini 3.6 FlashhighGoogle4.5%61.72026-09-16
23GPT-5.6 LunamaxOpenAI4.5%60.22026-09-16
24Gemini 3.5 FlashhighGoogle4.0%63.62026-09-16
25Qwen3.8 27BxhighAlibabaopen ↗4.0%59.22026-09-16
26DeepSeek V4 Pro 0813maxDeepSeekopen ↗1.0%56.02026-09-16
27Mercury 2.5highInception0.0%2026-09-16

Cite as: BenchLeader, “Terminal-Bench 4.0 (Vals) leaderboard”, https://www.benchleader.com/benchmarks/vals_terminal_bench_4, data as of 19 Sept 2026.

Terminal-Bench 4.0 (Vals): questions

What does Terminal-Bench 4.0 (Vals) measure?
Terminal-Bench 4.0 tasks written to defeat current frontier agents, run by Vals AI. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads Terminal-Bench 4.0 (Vals)?
GPT-6 Astra leads Terminal-Bench 4.0 (Vals) with 57.1% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 49.5%.
How many models have Terminal-Bench 4.0 (Vals) results?
27 model configurations have a Terminal-Bench 4.0 (Vals) result on BenchLeader, all taken from Vals AI.
Who runs Terminal-Bench 4.0 (Vals) and how often is it updated?
Terminal-Bench 4.0 (Vals) is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does Terminal-Bench 4.0 (Vals) count toward the BenchLeader Index?
No. Terminal-Bench 4.0 (Vals) is shown for reference but left out of the composite index.