BenchLeader

Terminal-Bench 2.0 (Vals)

Terminal-Bench 2.0 in Vals AI's fixed harness. Run by Vals AI.

As of 19 Sept 2026, GPT-5.5 leads Terminal-Bench 2.0 (Vals) on BenchLeader with 73.2%, ahead of Claude Opus 4.8 at 70.0%, across 64 model configurations with a published result.

Published by
Vals AI
Category
Agents & tools
Index weight
Reference only
Models
64
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

Terminal-Bench 2.0: agent tasks in a shell, checked by tests, run by Vals AI.

How it is scored

Percent of tasks passing, run by Vals AI.

What to keep in mind

Superseded by 2.1, kept for comparison.

64 of 64
#
1GPT-5.5highOpenAI73.2%67.02026-06-04
2Claude Opus 4.8Anthropic70.0%61.92026-06-04
3Claude Opus 4.7Anthropic68.5%64.52026-06-04
4Gemini 3.5 FlashhighGoogle67.4%63.62026-06-04
5Gemini 3.1 ProhighGoogle67.4%60.72026-06-04
6GPT-5.3 CodexhighOpenAI64.0%2026-06-04
7Muse SparkMeta59.5%65.72026-06-04
8Claude Sonnet 4.6Anthropic59.5%59.02026-06-04
9Qwen3 7maxAlibaba59.2%61.02026-06-04
10Claude Opus 4.6thinkingAnthropic58.4%62.52026-06-04
11GPT-5.4highOpenAI58.4%59.02026-06-04
12Claude Opus 4.5Anthropic58.4%58.22026-06-04
13Kimi K2.6Moonshot AIopen ↗57.3%60.52026-06-04
14DeepSeek V4 ProDeepSeekopen ↗56.2%52.62026-06-04
15Gemini 3 ProhighGoogle55.1%61.12026-06-04
16Claude Opus 4.5thinkingAnthropic53.9%60.82026-06-04
17GLM-5.1Zhipu AIopen ↗53.9%56.92026-06-04
18Qwen3 6maxAlibaba51.7%63.32026-06-04
19Gemini 3 FlashhighGoogle51.7%58.72026-06-04
20GPT-5.2highOpenAI51.7%55.72026-06-04
21GLM-5thinkingZhipu AIopen49.4%59.62026-06-04
22MiniMax-M2.7MiniMaxopen ↗47.2%56.12026-06-04
23MiniMax-M3MiniMaxopen ↗46.1%57.52026-06-04
24GPT-5.1highOpenAI44.9%58.72026-06-04
25Qwen3.6 PlusAlibabaopen44.9%58.52026-06-04
26GPT-5.4 minihighOpenAI44.9%53.02026-06-04
27Qwen3.6 27BAlibaba44.9%44.72026-06-04
28Grok 4.3xAI43.5%50.52026-06-04
29Qwen3.5 PlusthinkingAlibaba41.6%57.32026-06-04
30MiniMax-M2.5MiniMaxopen ↗41.6%54.22026-06-04
31Claude Sonnet 4.5thinkingAnthropic41.6%53.42026-06-04
32Grok 4.20thinkingxAI40.5%59.92026-06-04
33Kimi K2.5thinkingMoonshot AIopen40.5%58.52026-06-04
34GPT-5.4 nanohighOpenAI39.9%49.12026-06-04
35Gemma 4 31BhighGoogleopen ↗39.3%2026-06-04
36GLM-4.7Zhipu AIopen ↗38.2%52.42026-06-04
37Claude Haiku 4.5thinkingAnthropic38.2%47.52026-06-04
38GPT-5highOpenAI37.1%58.62026-06-04
39MiniMax-M2MiniMaxopen37.1%55.12026-06-04
40Kimi K2 ThinkingthinkingMoonshot AIopen37.1%52.22026-06-04
41DeepSeek V3.2highDeepSeekopen36.0%2026-06-04
42Laguna M.1Poolsideopen31.5%37.72026-06-04
43Gemini 2.5 ProGoogle30.3%54.52026-06-04
44Mistral Medium 3.5highMistral AIopen30.3%2026-06-04
45Grok 4 FastthinkingxAI29.2%52.62026-06-04
46Grok 4xAI28.1%58.52026-06-04
47GLM-4.6Zhipu AIopen28.1%49.82026-06-04
48Laguna XS.2Poolside28.1%36.52026-06-04
49GPT-5 minihighOpenAI27.0%55.32026-06-04
50Kimi K2Moonshot AIopen25.8%51.02026-06-04
51Grok 4.1thinkingxAI24.7%56.02026-06-04
52Qwen3 MaxmaxAlibaba24.7%53.52026-06-04
53Qwen3.5 FlashAlibaba24.7%52.52026-06-04
54Gemini 3.1 Flash LitehighGoogle24.7%49.12026-06-04
55Gemini 2.5 Flash 09 2025thinkingGoogle21.4%51.62026-06-04
56gpt-oss-120bOpenAIopen19.1%46.92026-06-04
57Trinity LargethinkingArcee AIopen18.0%47.62026-06-04
58Grok 4.1no reasoningxAI18.0%43.02026-06-04
59Mistral Small 4highMistral AIopen ↗16.9%2026-06-04
60GPT-4.1highOpenAI14.6%46.62026-06-04

Cite as: BenchLeader, “Terminal-Bench 2.0 (Vals) leaderboard”, https://www.benchleader.com/benchmarks/vals_terminal_bench_2, data as of 19 Sept 2026.

Terminal-Bench 2.0 (Vals): questions

What does Terminal-Bench 2.0 (Vals) measure?
Terminal-Bench 2.0: agent tasks in a shell, checked by tests, run by Vals AI. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads Terminal-Bench 2.0 (Vals)?
GPT-5.5 leads Terminal-Bench 2.0 (Vals) with 73.2% as of 19 Sept 2026, ahead of Claude Opus 4.8 at 70.0%.
How many models have Terminal-Bench 2.0 (Vals) results?
64 model configurations have a Terminal-Bench 2.0 (Vals) result on BenchLeader, all taken from Vals AI.
Who runs Terminal-Bench 2.0 (Vals) and how often is it updated?
Terminal-Bench 2.0 (Vals) is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does Terminal-Bench 2.0 (Vals) count toward the BenchLeader Index?
No. Terminal-Bench 2.0 (Vals) is shown for reference but left out of the composite index.