AutomationBench
Office automation tasks run end to end. Scored on partial credit, so a model that gets most of a workflow right still registers.
As of 22 Sept 2026, Claude Opus 5.5 leads AutomationBench on BenchLeader with 69.5%, ahead of DeepSeek V4.1 Flash at 68.9%, across 53 model configurations with a published result.
- Published by
- Artificial Analysis
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 53
- Data as of
- 22 Sept 2026
Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.
- 1Claude Opus 5.5 (thinking)69.5%
- 2DeepSeek V4.1 Flash (max)68.9%
- 3GPT-6 Astra (max)68.5%
- 4GPT-6 Astra (xhigh)67.2%
- 5Grok 4.6 (xhigh)67.0%
- 6Grok 4.6 (high)66.7%
- 7GPT-6 Astra (high)66.6%
- 8Grok 4.7 (xhigh)65.6%
- 9Claude Opus 5.5 (xhigh)65.0%
- 10GPT-6 Astra (medium)64.6%
- 11Grok 4.7 (high)63.5%
- 12Claude Opus 5.5 (high)63.2%
- 13GLM 5.3 (max)62.2%
- 14GLM 5.3 Flash60.4%
- 15GPT-5.6 Sol (max)60.1%
53 of 53
| # | ||||
|---|---|---|---|---|
| 1 | New | 69.5% | 70.7 | 2026-09-22 |
| 2 | 68.9% | 61.4 | 2026-09-10 | |
| 3 | 68.5% | 71.7 | 2026-09-03 | |
| 4 | 67.2% | 70.5 | 2026-09-03 | |
| 5 | 67.0% | 64.1 | 2026-08-12 | |
| 6 | 66.7% | 64.2 | 2026-08-12 | |
| 7 | 66.6% | 71.3 | 2026-09-03 | |
| 8 | New | 65.6% | 61.9 | 2026-09-21 |
| 9 | New | 65.0% | 70.3 | 2026-09-17 |
| 10 | 64.6% | 69.2 | 2026-09-03 | |
| 11 | New | 63.5% | – | 2026-09-21 |
| 12 | New | 63.2% | 69.9 | 2026-09-17 |
| 13 | 62.2% | 65.4 | 2026-08-18 | |
| 14 | 60.4% | 63.6 | 2026-08-26 | |
| 15 | 60.1% | 68.5 | 2026-07-09 | |
| 16 | 59.9% | 64.3 | 2026-09-02 | |
| 17 | 59.6% | 64.8 | 2026-07-09 | |
| 18 | 59.4% | 71.0 | 2026-09-01 | |
| 19 | 59.1% | 67.6 | 2026-09-03 | |
| 20 | New | 58.6% | – | 2026-09-21 |
| 21 | 58.3% | 66.7 | 2026-07-16 | |
| 22 | 57.9% | 69.2 | 2026-09-02 | |
| 23 | 57.8% | 71.1 | 2026-09-01 | |
| 24 | 57.3% | 64.5 | 2026-08-12 | |
| 25 | 56.8% | 67.9 | 2026-09-02 | |
| 26 | 56.7% | 63.7 | 2026-08-13 | |
| 27 | 56.6% | 69.5 | 2026-07-24 | |
| 28 | 56.2% | 66.2 | 2026-09-02 | |
| 29 | 55.3% | 71.0 | 2026-09-01 | |
| 30 | 54.7% | 69.1 | 2026-09-01 | |
| 31 | 54.3% | 66.4 | 2026-07-24 | |
| 32 | 54.1% | 70.3 | 2026-06-09 | |
| 33 | 53.6% | 69.8 | 2026-07-24 | |
| 34 | 53.2% | 69.3 | 2026-07-24 | |
| 35 | New | 51.0% | 64.8 | 2026-09-18 |
| 36 | 50.2% | 59.9 | 2026-07-09 | |
| 37 | 48.2% | 58.9 | 2026-08-14 | |
| 38 | 42.1% | 63.4 | 2026-05-19 | |
| 39 | 40.6% | 63.7 | 2026-08-05 | |
| 40 | 37.2% | 57.2 | 2026-09-03 | |
| 41 | 36.5% | 60.0 | 2026-06-30 | |
| 42 | 25.0% | 55.2 | 2026-07-21 | |
| 43 | 21.3% | 57.2 | 2026-06-01 | |
| 44 | 6.8% | 52.5 | 2026-08-10 | |
| 45 | 6.3% | 51.1 | 2026-04-29 | |
| 46 | 3.8% | 51.2 | 2026-03-11 | |
| 47 | 3.0% | 57.3 | 2026-06-04 | |
| 48 | 2.1% | 46.7 | 2026-08-25 | |
| 49 | 1.5% | 49.7 | 2026-08-25 | |
| 50 | 1.0% | 47.0 | 2025-12-15 | |
| 51 | 0.8% | 48.0 | 2026-08-11 | |
| 52 | 0.4% | 45.5 | 2026-08-25 | |
| 53 | 0.2% | 49.2 | 2025-08-05 |
Cite as: BenchLeader, “AutomationBench leaderboard”, https://www.benchleader.com/benchmarks/aa_automationbench, data as of 22 Sept 2026.
AutomationBench: questions
- What does AutomationBench measure?
- Office automation tasks run end to end. Scored on partial credit, so a model that gets most of a workflow right still registers. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads AutomationBench?
- Claude Opus 5.5 leads AutomationBench with 69.5% as of 22 Sept 2026, ahead of DeepSeek V4.1 Flash at 68.9%.
- How many models have AutomationBench results?
- 53 model configurations have a AutomationBench result on BenchLeader, all taken from Artificial Analysis.
- Who runs AutomationBench and how often is it updated?
- AutomationBench is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does AutomationBench count toward the BenchLeader Index?
- No. AutomationBench is shown for reference but left out of the composite index.