Terminal-Bench 2.0 (Vals)
Terminal-Bench 2.0 in Vals AI's fixed harness. Run by Vals AI.
As of 19 Sept 2026, GPT-5.5 leads Terminal-Bench 2.0 (Vals) on BenchLeader with 73.2%, ahead of Claude Opus 4.8 at 70.0%, across 64 model configurations with a published result.
- Published by
- Vals AI
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 64
- Data as of
- 19 Sept 2026
Vals AI (vals.ai).
What the test looks like
Terminal-Bench 2.0: agent tasks in a shell, checked by tests, run by Vals AI.
How it is scored
Percent of tasks passing, run by Vals AI.
What to keep in mind
Superseded by 2.1, kept for comparison.
- 1GPT-5.5 (high)73.2%
- 2Claude Opus 4.870.0%
- 3Claude Opus 4.768.5%
- 4Gemini 3.5 Flash (high)67.4%
- 5Gemini 3.1 Pro (high)67.4%
- 6GPT-5.3 Codex (high)64.0%
- 7Muse Spark59.5%
- 8Claude Sonnet 4.659.5%
- 9Qwen3 7 (max)59.2%
- 10Claude Opus 4.6 (thinking)58.4%
- 11GPT-5.4 (high)58.4%
- 12Claude Opus 4.558.4%
- 13Kimi K2.657.3%
- 14DeepSeek V4 Pro56.2%
- 15Gemini 3 Pro (high)55.1%
64 of 64
| # | ||||
|---|---|---|---|---|
| 1 | 73.2% | 67.0 | 2026-06-04 | |
| 2 | 70.0% | 61.9 | 2026-06-04 | |
| 3 | 68.5% | 64.5 | 2026-06-04 | |
| 4 | 67.4% | 63.6 | 2026-06-04 | |
| 5 | 67.4% | 60.7 | 2026-06-04 | |
| 6 | 64.0% | – | 2026-06-04 | |
| 7 | 59.5% | 65.7 | 2026-06-04 | |
| 8 | 59.5% | 59.0 | 2026-06-04 | |
| 9 | 59.2% | 61.0 | 2026-06-04 | |
| 10 | 58.4% | 62.5 | 2026-06-04 | |
| 11 | 58.4% | 59.0 | 2026-06-04 | |
| 12 | 58.4% | 58.2 | 2026-06-04 | |
| 13 | 57.3% | 60.5 | 2026-06-04 | |
| 14 | 56.2% | 52.6 | 2026-06-04 | |
| 15 | 55.1% | 61.1 | 2026-06-04 | |
| 16 | 53.9% | 60.8 | 2026-06-04 | |
| 17 | 53.9% | 56.9 | 2026-06-04 | |
| 18 | 51.7% | 63.3 | 2026-06-04 | |
| 19 | 51.7% | 58.7 | 2026-06-04 | |
| 20 | 51.7% | 55.7 | 2026-06-04 | |
| 21 | 49.4% | 59.6 | 2026-06-04 | |
| 22 | 47.2% | 56.1 | 2026-06-04 | |
| 23 | 46.1% | 57.5 | 2026-06-04 | |
| 24 | 44.9% | 58.7 | 2026-06-04 | |
| 25 | 44.9% | 58.5 | 2026-06-04 | |
| 26 | 44.9% | 53.0 | 2026-06-04 | |
| 27 | 44.9% | 44.7 | 2026-06-04 | |
| 28 | 43.5% | 50.5 | 2026-06-04 | |
| 29 | 41.6% | 57.3 | 2026-06-04 | |
| 30 | 41.6% | 54.2 | 2026-06-04 | |
| 31 | 41.6% | 53.4 | 2026-06-04 | |
| 32 | 40.5% | 59.9 | 2026-06-04 | |
| 33 | 40.5% | 58.5 | 2026-06-04 | |
| 34 | 39.9% | 49.1 | 2026-06-04 | |
| 35 | 39.3% | – | 2026-06-04 | |
| 36 | 38.2% | 52.4 | 2026-06-04 | |
| 37 | 38.2% | 47.5 | 2026-06-04 | |
| 38 | 37.1% | 58.6 | 2026-06-04 | |
| 39 | 37.1% | 55.1 | 2026-06-04 | |
| 40 | 37.1% | 52.2 | 2026-06-04 | |
| 41 | 36.0% | – | 2026-06-04 | |
| 42 | 31.5% | 37.7 | 2026-06-04 | |
| 43 | 30.3% | 54.5 | 2026-06-04 | |
| 44 | 30.3% | – | 2026-06-04 | |
| 45 | 29.2% | 52.6 | 2026-06-04 | |
| 46 | 28.1% | 58.5 | 2026-06-04 | |
| 47 | 28.1% | 49.8 | 2026-06-04 | |
| 48 | 28.1% | 36.5 | 2026-06-04 | |
| 49 | 27.0% | 55.3 | 2026-06-04 | |
| 50 | 25.8% | 51.0 | 2026-06-04 | |
| 51 | 24.7% | 56.0 | 2026-06-04 | |
| 52 | 24.7% | 53.5 | 2026-06-04 | |
| 53 | 24.7% | 52.5 | 2026-06-04 | |
| 54 | 24.7% | 49.1 | 2026-06-04 | |
| 55 | 21.4% | 51.6 | 2026-06-04 | |
| 56 | 19.1% | 46.9 | 2026-06-04 | |
| 57 | 18.0% | 47.6 | 2026-06-04 | |
| 58 | 18.0% | 43.0 | 2026-06-04 | |
| 59 | 16.9% | – | 2026-06-04 | |
| 60 | 14.6% | 46.6 | 2026-06-04 |
Cite as: BenchLeader, “Terminal-Bench 2.0 (Vals) leaderboard”, https://www.benchleader.com/benchmarks/vals_terminal_bench_2, data as of 19 Sept 2026.
Terminal-Bench 2.0 (Vals): questions
- What does Terminal-Bench 2.0 (Vals) measure?
- Terminal-Bench 2.0: agent tasks in a shell, checked by tests, run by Vals AI. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads Terminal-Bench 2.0 (Vals)?
- GPT-5.5 leads Terminal-Bench 2.0 (Vals) with 73.2% as of 19 Sept 2026, ahead of Claude Opus 4.8 at 70.0%.
- How many models have Terminal-Bench 2.0 (Vals) results?
- 64 model configurations have a Terminal-Bench 2.0 (Vals) result on BenchLeader, all taken from Vals AI.
- Who runs Terminal-Bench 2.0 (Vals) and how often is it updated?
- Terminal-Bench 2.0 (Vals) is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does Terminal-Bench 2.0 (Vals) count toward the BenchLeader Index?
- No. Terminal-Bench 2.0 (Vals) is shown for reference but left out of the composite index.