BenchLeader

GDPval (AA)

Artificial Analysis' normalised run of OpenAI's GDPval professional-task comparison.

As of 19 Sept 2026, Claude Fable 5.1 leads GDPval (AA) on BenchLeader with 62.3%, ahead of Claude Opus 5 at 61.8%, across 241 model configurations with a published result.

Published by
Artificial Analysis
Category
Agents & tools
Index weight
Reference only
Models
241
Data as of
19 Sept 2026

Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.

What the test looks like

GDPval asks models to produce real work products across 44 occupations and has professionals compare them with a human expert's. Artificial Analysis runs the public set and normalises the win rate.

How it is scored

Normalised win rate against the human deliverable, run by Artificial Analysis.

What to keep in mind

Human graders and pairwise comparison make it relative and slow to update; the Epoch-mirrored official results are counted in the index.

241 of 241
#
1Claude Fable 5.1xhighAnthropic62.3%71.6
2Claude Opus 5maxAnthropic61.8%69.9
3Claude Fable 5.1thinkingAnthropic61.2%71.3
4Claude Opus 5xhighAnthropic60.4%70.2
5Muse Spark 1.3maxMeta60.2%69.3
6Qwen3 8maxAlibaba58.2%66.5
7Grok 4.6xhighxAI58.1%64.6
8Muse Spark 1.3xhighMeta58.1%68.0
9GLM-5.3-FlashZhipu AIopen ↗57.8%63.9
10Claude Fable 5.1highAnthropic57.5%72.0
11Qwen3.8-Flash-NextAlibabaopen ↗57.4%61.6
12Grok 4.6highxAI57.1%64.5
13Grok 4.6mediumxAI57.1%66.2
14GLM-5.3maxZhipu AIopen ↗56.7%65.8
15DeepSeek V4.1 FlashmaxDeepSeekopen ↗56.6%61.9
16Claude Fable 5thinkingAnthropic56.6%70.9
17Claude Opus 5highAnthropic56.5%70.2
18Qwen3.8 2.4T A95BAlibabaopen ↗56.4%65.2
19GPT-5.6 SolmaxOpenAI54.3%68.8
20GPT-5.6 SolxhighOpenAI54.2%68.2
21GPT-6 AstramaxOpenAI54.0%71.8
22Claude Fable 5.1mediumAnthropic54.0%69.6
23Deepseek v4 Flash VisionmaxDeepSeekopen53.9%59.8
24Agnes 3.0 FlashSapiens AI53.7%62.3
25NewStep 5 PreviewStepFun53.6%65.0
26GPT-6 AstraxhighOpenAI52.8%71.1
27Kimi K3maxMoonshot AIopen ↗52.4%67.2
28GPT-6 AstrahighOpenAI51.5%71.9
29Claude Opus 5mediumAnthropic51.2%67.3
30GPT-5.6 SolhighOpenAI51.2%68.0
31Muse Spark 1.2xhighMeta51.2%64.1
32Claude Fable 5.1lowAnthropic50.2%67.5
33Claude Sonnet 5maxAnthropic50.0%60.5
34GPT-6 AstramediumOpenAI50.0%69.8
35DeepSeek V4 PromaxDeepSeekopen ↗49.7%64.2
36Claude Opus 4.8maxAnthropic49.5%64.2
37GPT-5.6 TerraxhighOpenAI48.9%64.3
38GPT-5.6 TerramaxOpenAI48.9%65.0
39Gemini 3.8 FlashhighGoogle48.2%64.5
40Qwen3.8 27BxhighAlibabaopen ↗48.2%59.2
41Grok 4.6lowxAI48.1%61.0
42GPT-5.6 LunamaxOpenAI47.8%60.2
43Gemini 3.8 FlashmediumGoogle47.8%64.9
44DeepSeek V4 FlashmaxDeepSeekopen ↗47.1%62.1
45Gemini 3.7 FlashhighGoogle46.7%64.6
46Qwen3.8 27BmediumAlibabaopen ↗46.5%54.6
47Grok 4.5highxAI46.5%61.0
48GPT-5.6 LunaxhighOpenAI46.4%61.1
49GPT-5.6 SolmediumOpenAI46.3%66.1
50GPT-6 AstralowOpenAI46.0%68.2
51GPT-5.6 TerrahighOpenAI45.8%62.8
52Gemini 3.7 FlashmediumGoogle45.3%64.7
53GLM-5.2maxZhipu AIopen ↗45.3%63.9
54Qwen3.8 27BlowAlibabaopen ↗45.1%53.4
55Claude Sonnet 5xhighAnthropic45.0%59.8
56K2 Horizon 375B A23BMBZUAIopen45.0%57.6
57GPT-5.5xhighOpenAI44.8%67.6
58Claude Opus 4.7maxAnthropic44.8%64.3
59Agnes 2.5 Pro BetaSapiens AI43.8%60.8
60GPT-5.6 LunahighOpenAI43.8%58.6

Cite as: BenchLeader, “GDPval (AA) leaderboard”, https://www.benchleader.com/benchmarks/aa_gdpval, data as of 19 Sept 2026.

GDPval (AA): questions

What does GDPval (AA) measure?
GDPval asks models to produce real work products across 44 occupations and has professionals compare them with a human expert's. Artificial Analysis runs the public set and normalises the win rate. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads GDPval (AA)?
Claude Fable 5.1 leads GDPval (AA) with 62.3% as of 19 Sept 2026, ahead of Claude Opus 5 at 61.8%.
How many models have GDPval (AA) results?
241 model configurations have a GDPval (AA) result on BenchLeader, all taken from Artificial Analysis.
Who runs GDPval (AA) and how often is it updated?
GDPval (AA) is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
Does GDPval (AA) count toward the BenchLeader Index?
No. GDPval (AA) is shown for reference but left out of the composite index.