Terminal-Bench 4.0 (Vals)
Terminal-Bench 4.0, the frontier-difficulty set, in Vals AI's harness. Run by Vals AI.
As of 19 Sept 2026, GPT-6 Astra leads Terminal-Bench 4.0 (Vals) on BenchLeader with 57.1%, ahead of Claude Fable 5.1 at 49.5%, across 27 model configurations with a published result.
- Published by
- Vals AI
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 27
- Data as of
- 19 Sept 2026
Vals AI (vals.ai).
What the test looks like
Terminal-Bench 4.0 tasks written to defeat current frontier agents, run by Vals AI.
How it is scored
Percent of tasks passing, run by Vals AI.
What to keep in mind
New and hard; small set, low scores.
- 1GPT-6 Astra (max)57.1%
- 2Claude Fable 5.149.5%
- 3Claude Opus 545.5%
- 4Muse Spark 1.3 (max)27.8%
- 5GPT-5.6 Sol (max)27.8%
- 6GPT-5.6 Terra (max)26.3%
- 7GLM-5.3 (max)25.3%
- 8Qwen3 8 (max)24.8%
- 9Claude Fable 522.7%
- 10GLM-5.3-Flash (max)19.7%
- 11Grok 4.6 (high)17.2%
- 12Claude Opus 4.816.2%
- 13Muse Spark 1.3 (xhigh)15.2%
- 14Gemini 3.8 Flash (high)13.1%
- 15Kimi K3 (max)12.6%
27 of 27
| # | ||||
|---|---|---|---|---|
| 1 | 57.1% | 71.8 | 2026-09-16 | |
| 2 | 49.5% | 64.5 | 2026-09-16 | |
| 3 | 45.5% | 66.7 | 2026-09-16 | |
| 4 | 27.8% | 69.3 | 2026-09-16 | |
| 5 | 27.8% | 68.8 | 2026-09-16 | |
| 6 | 26.3% | 65.0 | 2026-09-16 | |
| 7 | 25.3% | 65.8 | 2026-09-16 | |
| 8 | 24.8% | 66.5 | 2026-09-16 | |
| 9 | 22.7% | 68.3 | 2026-09-16 | |
| 10 | 19.7% | 55.3 | 2026-09-16 | |
| 11 | 17.2% | 64.5 | 2026-09-16 | |
| 12 | 16.2% | 61.9 | 2026-09-16 | |
| 13 | 15.2% | 68.0 | 2026-09-16 | |
| 14 | 13.1% | 64.5 | 2026-09-16 | |
| 15 | 12.6% | 67.2 | 2026-09-16 | |
| 16 | 11.6% | – | 2026-09-16 | |
| 17 | 9.1% | 55.0 | 2026-09-16 | |
| 18 | 8.1% | 57.2 | 2026-09-16 | |
| 19 | 6.6% | 61.0 | 2026-09-16 | |
| 20 | 6.1% | 64.6 | 2026-09-16 | |
| 21 | 5.6% | 64.1 | 2026-09-16 | |
| 22 | 4.5% | 61.7 | 2026-09-16 | |
| 23 | 4.5% | 60.2 | 2026-09-16 | |
| 24 | 4.0% | 63.6 | 2026-09-16 | |
| 25 | 4.0% | 59.2 | 2026-09-16 | |
| 26 | 1.0% | 56.0 | 2026-09-16 | |
| 27 | 0.0% | – | 2026-09-16 |
Cite as: BenchLeader, “Terminal-Bench 4.0 (Vals) leaderboard”, https://www.benchleader.com/benchmarks/vals_terminal_bench_4, data as of 19 Sept 2026.
Terminal-Bench 4.0 (Vals): questions
- What does Terminal-Bench 4.0 (Vals) measure?
- Terminal-Bench 4.0 tasks written to defeat current frontier agents, run by Vals AI. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads Terminal-Bench 4.0 (Vals)?
- GPT-6 Astra leads Terminal-Bench 4.0 (Vals) with 57.1% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 49.5%.
- How many models have Terminal-Bench 4.0 (Vals) results?
- 27 model configurations have a Terminal-Bench 4.0 (Vals) result on BenchLeader, all taken from Vals AI.
- Who runs Terminal-Bench 4.0 (Vals) and how often is it updated?
- Terminal-Bench 4.0 (Vals) is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does Terminal-Bench 4.0 (Vals) count toward the BenchLeader Index?
- No. Terminal-Bench 4.0 (Vals) is shown for reference but left out of the composite index.