BenchLeader

GDP.pdf

Document work on real PDFs, scored all-or-nothing: every required element has to be right for the task to pass, which is why the scores look low.

As of 22 Sept 2026, GPT-6 Astra leads GDP.pdf on BenchLeader with 32.2%, ahead of GPT-6 Astra at 31.0%, across 53 model configurations with a published result.

Published by
Artificial Analysis
Category
Agents & tools
Index weight
Reference only
Models
53
Data as of
22 Sept 2026

Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.

53 of 53
#
1GPT-6 AstraxhighOpenAI32.2%70.52026-09-03
2GPT-6 AstramaxOpenAI31.0%71.72026-09-03
3GPT-6 AstrahighOpenAI31.0%71.32026-09-03
4GPT-6 AstramediumOpenAI30.4%69.22026-09-03
5GPT-6 AstralowOpenAI30.4%67.62026-09-03
6NewClaude Opus 5.5highAnthropic28.8%69.92026-09-17
7GPT-5.6 SolmaxOpenAI27.2%68.52026-07-09
8Claude Fable 5.1highAnthropic26.8%71.02026-09-01
9Claude Fable 5.1mediumAnthropic26.8%69.12026-09-01
10NewClaude Opus 5.5xhighAnthropic26.6%70.32026-09-17
11Muse Spark 1.3maxMeta26.6%69.22026-09-02
12Claude Fable 5.1xhighAnthropic26.2%71.12026-09-01
13Claude Fable 5.1thinkingAnthropic26.2%71.02026-09-01
14NewClaude Opus 5.5thinkingAnthropic26.2%70.72026-09-22
15Muse Spark 1.3xhighMeta24.2%67.92026-09-02
16Claude Fable 5thinkingAnthropic24.0%70.32026-06-09
17GPT-5.6 TerramaxOpenAI24.0%64.82026-07-09
18GPT-5.6 LunamaxOpenAI24.0%59.92026-07-09
19NewGrok 4.7highSpaceXAI23.2%2026-09-21
20Qwen3.8 Max (0902)maxAlibaba22.8%66.22026-09-02
21Kimi K3maxMoonshot AIopen ↗22.0%66.72026-07-16
22Claude Opus 5maxAnthropic21.6%69.52026-07-24
23Claude Opus 5xhighAnthropic21.0%69.32026-07-24
24Gemini 3.8 FlashhighGoogle21.0%64.32026-09-02
25Claude Opus 5mediumAnthropic20.0%66.42026-07-24
26NewGrok 4.7xhighSpaceXAI20.0%61.92026-09-21
27Gemini 3.5 FlashhighGoogle19.8%63.42026-05-19
28Claude Opus 5highAnthropic19.6%69.82026-07-24
29NewMiMo-V2.6-ProXiaomiopen19.2%2026-09-21
30Muse Spark 1.2xhighMeta17.4%63.72026-08-05
31Grok 4.6xhighSpaceXAI17.2%64.12026-08-12
32Grok 4.6highSpaceXAI17.0%64.22026-08-12
33Qwen3.8 27BxhighAlibabaopen ↗16.6%58.92026-08-14
34GLM 5.3 FlashZhipu AIopen ↗15.4%63.62026-08-26
35Qwen3.8 2.4T A95BAlibabaopen ↗15.0%64.52026-08-12
36NewStep 5 PreviewStepFun14.8%64.82026-09-18
37Gemini 3.5 Flash LiteGoogle13.6%55.22026-07-21
38Claude Sonnet 5maxAnthropic13.2%60.02026-06-30
39DeepSeek V4.1 FlashmaxDeepSeekopen ↗12.8%61.42026-09-10
40DeepSeek V4 PromaxDeepSeekopen ↗11.4%63.72026-08-13
41GLM 5.3maxZhipu AIopen ↗11.2%65.42026-08-18
42Muse GlimmerhighMetaopen ↗10.0%52.52026-08-10
43MiniMax M3MiniMaxopen ↗9.8%57.22026-06-01
44K2 Horizon 375B A23BMBZUAIopen7.4%57.22026-09-03
45Nemotron 3 Ultra 550B A55BthinkingNVIDIAopen ↗5.0%57.32026-06-04
46gpt-oss-120bhighOpenAIopen ↗4.0%49.22025-08-05
47Nemotron 3.5 LightningNVIDIAopen ↗3.0%48.02026-08-11
48Mistral Medium 3.5Mistral AIopen2.8%51.12026-04-29
49Nemotron 3 Super 120B A12bthinkingNVIDIAopen ↗2.6%51.22026-03-11
50Granite 4.2 30BIBMopen ↗1.8%49.72026-08-25
51Nemotron 3 Nano 30B A3BthinkingNVIDIAopen ↗1.0%47.02025-12-15
52Granite 4.2 8BIBMopen ↗0.6%46.72026-08-25
53Granite 4.2 3BIBMopen ↗0.2%45.52026-08-25

Cite as: BenchLeader, “GDP.pdf leaderboard”, https://www.benchleader.com/benchmarks/aa_gdp_pdf, data as of 22 Sept 2026.

GDP.pdf: questions

What does GDP.pdf measure?
Document work on real PDFs, scored all-or-nothing: every required element has to be right for the task to pass, which is why the scores look low. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads GDP.pdf?
GPT-6 Astra leads GDP.pdf with 32.2% as of 22 Sept 2026, ahead of GPT-6 Astra at 31.0%.
How many models have GDP.pdf results?
53 model configurations have a GDP.pdf result on BenchLeader, all taken from Artificial Analysis.
Who runs GDP.pdf and how often is it updated?
GDP.pdf is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
Does GDP.pdf count toward the BenchLeader Index?
No. GDP.pdf is shown for reference but left out of the composite index.