AA-Briefcase
Knowledge work delivered as real files — spreadsheets, decks, documents — graded head to head, so the score is an Elo rather than a percentage.
As of 22 Sept 2026, Claude Opus 5.5 leads AA-Briefcase on BenchLeader with 1822, ahead of Claude Opus 5.5 at 1780, across 51 model configurations with a published result.
- Published by
- Artificial Analysis
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 51
- Data as of
- 22 Sept 2026
Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.
- 1Claude Opus 5.5 (thinking)1822
- 2Claude Opus 5.5 (xhigh)1780
- 3Claude Opus 5.5 (high)1705
- 4Claude Fable 5.1 (thinking)1678
- 5Claude Opus 5 (max)1673
- 6Claude Fable 5.1 (xhigh)1669
- 7Grok 4.7 (xhigh)1657
- 8Claude Opus 5 (xhigh)1649
- 9Grok 4.7 (high)1644
- 10Qwen3.8 Max (0902) (max)1640
- 11Muse Spark 1.3 (max)1597
- 12Claude Fable 5.1 (high)1592
- 13Claude Opus 5 (high)1573
- 14GPT-6 Astra (max)1569
- 15Grok 4.6 (xhigh)1555
51 of 51
| # | ||||
|---|---|---|---|---|
| 1 | New | 1822 | 70.7 | 2026-09-22 |
| 2 | New | 1780 | 70.3 | 2026-09-17 |
| 3 | New | 1705 | 69.9 | 2026-09-17 |
| 4 | 1678 | 71.0 | 2026-09-01 | |
| 5 | 1673 | 69.5 | 2026-07-24 | |
| 6 | 1669 | 71.1 | 2026-09-01 | |
| 7 | New | 1657 | 61.9 | 2026-09-21 |
| 8 | 1649 | 69.3 | 2026-07-24 | |
| 9 | New | 1644 | – | 2026-09-21 |
| 10 | 1640 | 66.2 | 2026-09-02 | |
| 11 | 1597 | 69.2 | 2026-09-02 | |
| 12 | 1592 | 71.0 | 2026-09-01 | |
| 13 | 1573 | 69.8 | 2026-07-24 | |
| 14 | 1569 | 71.7 | 2026-09-03 | |
| 15 | 1555 | 64.1 | 2026-08-12 | |
| 16 | 1546 | 64.2 | 2026-08-12 | |
| 17 | 1544 | 70.5 | 2026-09-03 | |
| 18 | 1543 | 70.3 | 2026-06-09 | |
| 19 | 1542 | 69.1 | 2026-09-01 | |
| 20 | 1525 | 65.4 | 2026-08-18 | |
| 21 | New | 1522 | – | 2026-09-21 |
| 22 | 1510 | 66.7 | 2026-07-16 | |
| 23 | 1507 | 71.3 | 2026-09-03 | |
| 24 | 1494 | 67.9 | 2026-09-02 | |
| 25 | 1487 | 68.5 | 2026-07-09 | |
| 26 | 1459 | 63.6 | 2026-08-26 | |
| 27 | 1459 | 69.2 | 2026-09-03 | |
| 28 | 1445 | 64.5 | 2026-08-12 | |
| 29 | 1441 | 66.4 | 2026-07-24 | |
| 30 | New | 1432 | 64.8 | 2026-09-18 |
| 31 | 1432 | 61.4 | 2026-09-10 | |
| 32 | 1403 | 58.9 | 2026-08-14 | |
| 33 | 1363 | 60.0 | 2026-06-30 | |
| 34 | 1345 | 59.9 | 2026-07-09 | |
| 35 | 1340 | 63.7 | 2026-08-05 | |
| 36 | 1336 | 64.8 | 2026-07-09 | |
| 37 | 1303 | 57.2 | 2026-09-03 | |
| 38 | 1261 | 63.7 | 2026-08-13 | |
| 39 | 1261 | 67.6 | 2026-09-03 | |
| 40 | 1202 | 64.3 | 2026-09-02 | |
| 41 | 1092 | 57.2 | 2026-06-01 | |
| 42 | 873 | 57.3 | 2026-06-04 | |
| 43 | 868 | 63.4 | 2026-05-19 | |
| 44 | 642 | 55.2 | 2026-07-21 | |
| 45 | 515 | 51.1 | 2026-04-29 | |
| 46 | 500 | 48.0 | 2026-08-11 | |
| 47 | 469 | 52.5 | 2026-08-10 | |
| 48 | 302 | 46.7 | 2026-08-25 | |
| 49 | 94 | 45.5 | 2026-08-25 | |
| 50 | 75 | 47.0 | 2025-12-15 | |
| 51 | 4 | 48.4 | 2025-12-15 |
Cite as: BenchLeader, “AA-Briefcase leaderboard”, https://www.benchleader.com/benchmarks/aa_briefcase, data as of 22 Sept 2026.
AA-Briefcase: questions
- What does AA-Briefcase measure?
- Knowledge work delivered as real files — spreadsheets, decks, documents — graded head to head, so the score is an Elo rather than a percentage. Scores are reported in Elo-style ratings from pairwise comparisons; higher is better.
- Which AI model leads AA-Briefcase?
- Claude Opus 5.5 leads AA-Briefcase with 1822 as of 22 Sept 2026, ahead of Claude Opus 5.5 at 1780.
- How many models have AA-Briefcase results?
- 51 model configurations have a AA-Briefcase result on BenchLeader, all taken from Artificial Analysis.
- Who runs AA-Briefcase and how often is it updated?
- AA-Briefcase is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does AA-Briefcase count toward the BenchLeader Index?
- No. AA-Briefcase is shown for reference but left out of the composite index.