BenchLeader

AA-Briefcase

Knowledge work delivered as real files — spreadsheets, decks, documents — graded head to head, so the score is an Elo rather than a percentage.

As of 22 Sept 2026, Claude Opus 5.5 leads AA-Briefcase on BenchLeader with 1822, ahead of Claude Opus 5.5 at 1780, across 51 model configurations with a published result.

Published by
Artificial Analysis
Category
Agents & tools
Index weight
Reference only
Models
51
Data as of
22 Sept 2026

Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.

51 of 51
#
1NewClaude Opus 5.5thinkingAnthropic182270.72026-09-22
2NewClaude Opus 5.5xhighAnthropic178070.32026-09-17
3NewClaude Opus 5.5highAnthropic170569.92026-09-17
4Claude Fable 5.1thinkingAnthropic167871.02026-09-01
5Claude Opus 5maxAnthropic167369.52026-07-24
6Claude Fable 5.1xhighAnthropic166971.12026-09-01
7NewGrok 4.7xhighSpaceXAI165761.92026-09-21
8Claude Opus 5xhighAnthropic164969.32026-07-24
9NewGrok 4.7highSpaceXAI16442026-09-21
10Qwen3.8 Max (0902)maxAlibaba164066.22026-09-02
11Muse Spark 1.3maxMeta159769.22026-09-02
12Claude Fable 5.1highAnthropic159271.02026-09-01
13Claude Opus 5highAnthropic157369.82026-07-24
14GPT-6 AstramaxOpenAI156971.72026-09-03
15Grok 4.6xhighSpaceXAI155564.12026-08-12
16Grok 4.6highSpaceXAI154664.22026-08-12
17GPT-6 AstraxhighOpenAI154470.52026-09-03
18Claude Fable 5thinkingAnthropic154370.32026-06-09
19Claude Fable 5.1mediumAnthropic154269.12026-09-01
20GLM 5.3maxZhipu AIopen ↗152565.42026-08-18
21NewMiMo-V2.6-ProXiaomiopen15222026-09-21
22Kimi K3maxMoonshot AIopen ↗151066.72026-07-16
23GPT-6 AstrahighOpenAI150771.32026-09-03
24Muse Spark 1.3xhighMeta149467.92026-09-02
25GPT-5.6 SolmaxOpenAI148768.52026-07-09
26GLM 5.3 FlashZhipu AIopen ↗145963.62026-08-26
27GPT-6 AstramediumOpenAI145969.22026-09-03
28Qwen3.8 2.4T A95BAlibabaopen ↗144564.52026-08-12
29Claude Opus 5mediumAnthropic144166.42026-07-24
30NewStep 5 PreviewStepFun143264.82026-09-18
31DeepSeek V4.1 FlashmaxDeepSeekopen ↗143261.42026-09-10
32Qwen3.8 27BxhighAlibabaopen ↗140358.92026-08-14
33Claude Sonnet 5maxAnthropic136360.02026-06-30
34GPT-5.6 LunamaxOpenAI134559.92026-07-09
35Muse Spark 1.2xhighMeta134063.72026-08-05
36GPT-5.6 TerramaxOpenAI133664.82026-07-09
37K2 Horizon 375B A23BMBZUAIopen130357.22026-09-03
38DeepSeek V4 PromaxDeepSeekopen ↗126163.72026-08-13
39GPT-6 AstralowOpenAI126167.62026-09-03
40Gemini 3.8 FlashhighGoogle120264.32026-09-02
41MiniMax M3MiniMaxopen ↗109257.22026-06-01
42Nemotron 3 Ultra 550B A55BthinkingNVIDIAopen ↗87357.32026-06-04
43Gemini 3.5 FlashhighGoogle86863.42026-05-19
44Gemini 3.5 Flash LiteGoogle64255.22026-07-21
45Mistral Medium 3.5Mistral AIopen51551.12026-04-29
46Nemotron 3.5 LightningNVIDIAopen ↗50048.02026-08-11
47Muse GlimmerhighMetaopen ↗46952.52026-08-10
48Granite 4.2 8BIBMopen ↗30246.72026-08-25
49Granite 4.2 3BIBMopen ↗9445.52026-08-25
50Nemotron 3 Nano 30B A3BthinkingNVIDIAopen ↗7547.02025-12-15
51K2 Think V2thinkingMBZUAIopen448.42025-12-15

Cite as: BenchLeader, “AA-Briefcase leaderboard”, https://www.benchleader.com/benchmarks/aa_briefcase, data as of 22 Sept 2026.

AA-Briefcase: questions

What does AA-Briefcase measure?
Knowledge work delivered as real files — spreadsheets, decks, documents — graded head to head, so the score is an Elo rather than a percentage. Scores are reported in Elo-style ratings from pairwise comparisons; higher is better.
Which AI model leads AA-Briefcase?
Claude Opus 5.5 leads AA-Briefcase with 1822 as of 22 Sept 2026, ahead of Claude Opus 5.5 at 1780.
How many models have AA-Briefcase results?
51 model configurations have a AA-Briefcase result on BenchLeader, all taken from Artificial Analysis.
Who runs AA-Briefcase and how often is it updated?
AA-Briefcase is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
Does AA-Briefcase count toward the BenchLeader Index?
No. AA-Briefcase is shown for reference but left out of the composite index.