BenchLeader

AutomationBench

Office automation tasks run end to end. Scored on partial credit, so a model that gets most of a workflow right still registers.

As of 22 Sept 2026, Claude Opus 5.5 leads AutomationBench on BenchLeader with 69.5%, ahead of DeepSeek V4.1 Flash at 68.9%, across 53 model configurations with a published result.

Published by
Artificial Analysis
Category
Agents & tools
Index weight
Reference only
Models
53
Data as of
22 Sept 2026

Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.

53 of 53
#
1NewClaude Opus 5.5thinkingAnthropic69.5%70.72026-09-22
2DeepSeek V4.1 FlashmaxDeepSeekopen ↗68.9%61.42026-09-10
3GPT-6 AstramaxOpenAI68.5%71.72026-09-03
4GPT-6 AstraxhighOpenAI67.2%70.52026-09-03
5Grok 4.6xhighSpaceXAI67.0%64.12026-08-12
6Grok 4.6highSpaceXAI66.7%64.22026-08-12
7GPT-6 AstrahighOpenAI66.6%71.32026-09-03
8NewGrok 4.7xhighSpaceXAI65.6%61.92026-09-21
9NewClaude Opus 5.5xhighAnthropic65.0%70.32026-09-17
10GPT-6 AstramediumOpenAI64.6%69.22026-09-03
11NewGrok 4.7highSpaceXAI63.5%2026-09-21
12NewClaude Opus 5.5highAnthropic63.2%69.92026-09-17
13GLM 5.3maxZhipu AIopen ↗62.2%65.42026-08-18
14GLM 5.3 FlashZhipu AIopen ↗60.4%63.62026-08-26
15GPT-5.6 SolmaxOpenAI60.1%68.52026-07-09
16Gemini 3.8 FlashhighGoogle59.9%64.32026-09-02
17GPT-5.6 TerramaxOpenAI59.6%64.82026-07-09
18Claude Fable 5.1thinkingAnthropic59.4%71.02026-09-01
19GPT-6 AstralowOpenAI59.1%67.62026-09-03
20NewMiMo-V2.6-ProXiaomiopen58.6%2026-09-21
21Kimi K3maxMoonshot AIopen ↗58.3%66.72026-07-16
22Muse Spark 1.3maxMeta57.9%69.22026-09-02
23Claude Fable 5.1xhighAnthropic57.8%71.12026-09-01
24Qwen3.8 2.4T A95BAlibabaopen ↗57.3%64.52026-08-12
25Muse Spark 1.3xhighMeta56.8%67.92026-09-02
26DeepSeek V4 PromaxDeepSeekopen ↗56.7%63.72026-08-13
27Claude Opus 5maxAnthropic56.6%69.52026-07-24
28Qwen3.8 Max (0902)maxAlibaba56.2%66.22026-09-02
29Claude Fable 5.1highAnthropic55.3%71.02026-09-01
30Claude Fable 5.1mediumAnthropic54.7%69.12026-09-01
31Claude Opus 5mediumAnthropic54.3%66.42026-07-24
32Claude Fable 5thinkingAnthropic54.1%70.32026-06-09
33Claude Opus 5highAnthropic53.6%69.82026-07-24
34Claude Opus 5xhighAnthropic53.2%69.32026-07-24
35NewStep 5 PreviewStepFun51.0%64.82026-09-18
36GPT-5.6 LunamaxOpenAI50.2%59.92026-07-09
37Qwen3.8 27BxhighAlibabaopen ↗48.2%58.92026-08-14
38Gemini 3.5 FlashhighGoogle42.1%63.42026-05-19
39Muse Spark 1.2xhighMeta40.6%63.72026-08-05
40K2 Horizon 375B A23BMBZUAIopen37.2%57.22026-09-03
41Claude Sonnet 5maxAnthropic36.5%60.02026-06-30
42Gemini 3.5 Flash LiteGoogle25.0%55.22026-07-21
43MiniMax M3MiniMaxopen ↗21.3%57.22026-06-01
44Muse GlimmerhighMetaopen ↗6.8%52.52026-08-10
45Mistral Medium 3.5Mistral AIopen6.3%51.12026-04-29
46Nemotron 3 Super 120B A12bthinkingNVIDIAopen ↗3.8%51.22026-03-11
47Nemotron 3 Ultra 550B A55BthinkingNVIDIAopen ↗3.0%57.32026-06-04
48Granite 4.2 8BIBMopen ↗2.1%46.72026-08-25
49Granite 4.2 30BIBMopen ↗1.5%49.72026-08-25
50Nemotron 3 Nano 30B A3BthinkingNVIDIAopen ↗1.0%47.02025-12-15
51Nemotron 3.5 LightningNVIDIAopen ↗0.8%48.02026-08-11
52Granite 4.2 3BIBMopen ↗0.4%45.52026-08-25
53gpt-oss-120bhighOpenAIopen ↗0.2%49.22025-08-05

Cite as: BenchLeader, “AutomationBench leaderboard”, https://www.benchleader.com/benchmarks/aa_automationbench, data as of 22 Sept 2026.

AutomationBench: questions

What does AutomationBench measure?
Office automation tasks run end to end. Scored on partial credit, so a model that gets most of a workflow right still registers. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads AutomationBench?
Claude Opus 5.5 leads AutomationBench with 69.5% as of 22 Sept 2026, ahead of DeepSeek V4.1 Flash at 68.9%.
How many models have AutomationBench results?
53 model configurations have a AutomationBench result on BenchLeader, all taken from Artificial Analysis.
Who runs AutomationBench and how often is it updated?
AutomationBench is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
Does AutomationBench count toward the BenchLeader Index?
No. AutomationBench is shown for reference but left out of the composite index.