Terminal-Bench 2.1 (AA)
The previous Terminal-Bench release, kept because far more models have been run on it than on 4.0.
As of 22 Sept 2026, Claude Fable 5.1 leads Terminal-Bench 2.1 (AA) on BenchLeader with 91.4%, ahead of Claude Fable 5.1 at 91.0%, across 236 model configurations with a published result.
- Published by
- Artificial Analysis
- Category
- Coding
- Index weight
- Reference only
- Models
- 236
- Data as of
- 22 Sept 2026
Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.
- 1Claude Fable 5.1 (thinking)91.4%
- 2Claude Fable 5.1 (xhigh)91.0%
- 3GPT-6 Astra (high)89.9%
- 4Claude Fable 5.1 (high)89.9%
- 5GPT-6 Astra (medium)89.5%
- 6GPT-5.6 Sol (xhigh)89.5%
- 7GPT-6 Astra (xhigh)89.1%
- 8Claude Opus 5 (max)89.1%
- 9Qwen3.8 Max (0902) (max)88.8%
- 10GPT-6 Astra (max)88.4%
- 11Grok 4.6 (high)88.4%
- 12Claude Opus 5 (xhigh)88.0%
- 13Claude Fable 5.1 (medium)88.0%
- 14GPT-5.6 Sol (max)88.0%
- 15GPT-6 Astra (low)88.0%
236 of 236
| # | |||
|---|---|---|---|
| 1 | 91.4% | 71.0 | |
| 2 | 91.0% | 71.1 | |
| 3 | 89.9% | 71.3 | |
| 4 | 89.9% | 71.0 | |
| 5 | 89.5% | 69.2 | |
| 6 | 89.5% | 67.9 | |
| 7 | 89.1% | 70.5 | |
| 8 | 89.1% | 69.5 | |
| 9 | 88.8% | 66.2 | |
| 10 | 88.4% | 71.7 | |
| 11 | 88.4% | 64.2 | |
| 12 | 88.0% | 69.3 | |
| 13 | 88.0% | 69.1 | |
| 14 | 88.0% | 68.5 | |
| 15 | 88.0% | 67.6 | |
| 16 | 88.0% | 64.8 | |
| 17 | 88.0% | 64.1 | |
| 18 | 87.6% | 69.8 | |
| 19 | 87.6% | 64.3 | |
| 20 | 87.3% | 67.4 | |
| 21 | 86.1% | 66.4 | |
| 22 | 86.1% | 65.5 | |
| 23 | 86.1% | 61.2 | |
| 24 | 85.8% | 64.3 | |
| 25 | 85.4% | 67.9 | |
| 26 | 85.0% | 66.7 | |
| 27 | 85.0% | 66.7 | |
| 28 | 84.6% | 70.3 | |
| 29 | 84.6% | 63.5 | |
| 30 | 84.3% | 69.2 | |
| 31 | 84.3% | 67.2 | |
| 32 | 84.3% | 65.5 | |
| 33 | 84.3% | 63.6 | |
| 34 | 83.9% | 65.4 | |
| 35 | 83.9% | 64.6 | |
| 36 | 83.2% | 63.7 | |
| 37 | 83.2% | 60.2 | |
| 38 | 82.4% | 59.7 | |
| 39 | 82.0% | 64.5 | |
| 40 | 82.0% | 61.7 | |
| 41 | 81.7% | 60.6 | |
| 42 | 80.9% | 59.9 | |
| 43 | 80.5% | 64.0 | |
| 44 | 80.5% | 60.0 | |
| 45 | 80.2% | 63.9 | |
| 46 | 80.2% | 63.7 | |
| 47 | 79.8% | 61.6 | |
| 48 | 79.8% | 58.9 | |
| 49 | 79.4% | 66.6 | |
| 50 | 78.7% | 63.7 | |
| 51 | 78.7% | 63.4 | |
| 52 | 78.7% | 61.6 | |
| 53 | 78.3% | 65.1 | |
| 54 | 78.3% | 64.4 | |
| 55 | 77.9% | 63.6 | |
| 56 | 77.9% | 61.4 | |
| 57 | 77.9% | 60.8 | |
| 58 | 77.5% | 61.5 | |
| 59 | 76.8% | 63.2 | |
| 60 | 76.4% | 62.8 |
Cite as: BenchLeader, “Terminal-Bench 2.1 (AA) leaderboard”, https://www.benchleader.com/benchmarks/aa_terminalbench_21, data as of 22 Sept 2026.
Terminal-Bench 2.1 (AA): questions
- What does Terminal-Bench 2.1 (AA) measure?
- The previous Terminal-Bench release, kept because far more models have been run on it than on 4.0. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads Terminal-Bench 2.1 (AA)?
- Claude Fable 5.1 leads Terminal-Bench 2.1 (AA) with 91.4% as of 22 Sept 2026, ahead of Claude Fable 5.1 at 91.0%.
- How many models have Terminal-Bench 2.1 (AA) results?
- 236 model configurations have a Terminal-Bench 2.1 (AA) result on BenchLeader, all taken from Artificial Analysis.
- Who runs Terminal-Bench 2.1 (AA) and how often is it updated?
- Terminal-Bench 2.1 (AA) is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does Terminal-Bench 2.1 (AA) count toward the BenchLeader Index?
- No. Terminal-Bench 2.1 (AA) is shown for reference but left out of the composite index.