Terminal-Bench 4.0 (AA)
Real command-line tasks in a sandboxed terminal, version 4.0. Run by Artificial Analysis; the Terminal-Bench team's own run is the one in the index.
As of 22 Sept 2026, Claude Opus 5.5 leads Terminal-Bench 4.0 (AA) on BenchLeader with 59.6%, ahead of GPT-6 Astra at 59.6%, across 167 model configurations with a published result.
- Published by
- Artificial Analysis
- Category
- Coding
- Index weight
- Reference only
- Models
- 167
- Data as of
- 22 Sept 2026
Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.
- 1Claude Opus 5.5 (thinking)59.6%
- 2GPT-6 Astra (xhigh)59.6%
- 3Claude Opus 5.5 (xhigh)59.6%
- 4GPT-6 Astra (max)59.1%
- 5Claude Opus 5.5 (high)56.6%
- 6Claude Fable 5.1 (xhigh)55.0%
- 7GPT-6 Astra (high)54.0%
- 8Claude Opus 5.5 (medium)52.5%
- 9Claude Fable 5.1 (high)52.0%
- 10Claude Fable 5.1 (thinking)52.0%
- 11GPT-6 Astra (medium)49.5%
- 12Claude Opus 5 (max)49.0%
- 13Claude Opus 5 (xhigh)46.5%
- 14Claude Opus 5 (high)46.0%
- 15Claude Fable 5.1 (medium)45.0%
167 of 167
| # | |||
|---|---|---|---|
| 1 | New | 59.6% | 70.7 |
| 2 | 59.6% | 70.5 | |
| 3 | New | 59.6% | 70.3 |
| 4 | 59.1% | 71.7 | |
| 5 | New | 56.6% | 69.9 |
| 6 | 55.0% | 71.1 | |
| 7 | 54.0% | 71.3 | |
| 8 | New | 52.5% | 69.3 |
| 9 | 52.0% | 71.0 | |
| 10 | 52.0% | 71.0 | |
| 11 | 49.5% | 69.2 | |
| 12 | 49.0% | 69.5 | |
| 13 | 46.5% | 69.3 | |
| 14 | 46.0% | 69.8 | |
| 15 | 45.0% | 69.1 | |
| 16 | 42.4% | 70.3 | |
| 17 | 41.9% | 67.6 | |
| 18 | 41.9% | 65.4 | |
| 19 | 40.4% | 66.7 | |
| 20 | 39.9% | 68.5 | |
| 21 | 38.9% | 66.2 | |
| 22 | 35.4% | 64.8 | |
| 23 | New | 34.9% | – |
| 24 | 34.3% | 66.4 | |
| 25 | 33.3% | 69.2 | |
| 26 | New | 33.3% | 64.8 |
| 27 | 32.8% | 63.6 | |
| 28 | New | 31.3% | 65.2 |
| 29 | 26.8% | 61.4 | |
| 30 | 26.3% | 62.8 | |
| 31 | New | 25.8% | 61.9 |
| 32 | 25.3% | 61.2 | |
| 33 | 24.8% | 67.9 | |
| 34 | New | 24.8% | – |
| 35 | 21.7% | 63.5 | |
| 36 | 21.2% | 64.2 | |
| 37 | 20.7% | 67.4 | |
| 38 | 19.7% | 64.6 | |
| 39 | 19.7% | 64.3 | |
| 40 | 17.2% | 64.1 | |
| 41 | 16.7% | 67.9 | |
| 42 | 14.7% | 67.2 | |
| 43 | 14.7% | 65.5 | |
| 44 | 14.7% | 63.7 | |
| 45 | 14.1% | 60.0 | |
| 46 | 13.6% | 64.3 | |
| 47 | 13.1% | 65.5 | |
| 48 | 12.6% | 66.7 | |
| 49 | 12.6% | 59.7 | |
| 50 | 12.6% | 58.8 | |
| 51 | 12.1% | 61.6 | |
| 52 | 12.1% | 59.3 | |
| 53 | 11.6% | 59.9 | |
| 54 | 11.1% | 64.5 | |
| 55 | 10.6% | 60.6 | |
| 56 | 10.1% | 63.9 | |
| 57 | 10.1% | 60.2 | |
| 58 | 9.1% | 66.6 | |
| 59 | 7.1% | 63.7 | |
| 60 | 7.1% | 61.7 |
Cite as: BenchLeader, “Terminal-Bench 4.0 (AA) leaderboard”, https://www.benchleader.com/benchmarks/aa_terminalbench_40, data as of 22 Sept 2026.
Terminal-Bench 4.0 (AA): questions
- What does Terminal-Bench 4.0 (AA) measure?
- Real command-line tasks in a sandboxed terminal, version 4.0. Run by Artificial Analysis; the Terminal-Bench team's own run is the one in the index. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads Terminal-Bench 4.0 (AA)?
- Claude Opus 5.5 leads Terminal-Bench 4.0 (AA) with 59.6% as of 22 Sept 2026, ahead of GPT-6 Astra at 59.6%.
- How many models have Terminal-Bench 4.0 (AA) results?
- 167 model configurations have a Terminal-Bench 4.0 (AA) result on BenchLeader, all taken from Artificial Analysis.
- Who runs Terminal-Bench 4.0 (AA) and how often is it updated?
- Terminal-Bench 4.0 (AA) is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does Terminal-Bench 4.0 (AA) count toward the BenchLeader Index?
- No. Terminal-Bench 4.0 (AA) is shown for reference but left out of the composite index.