BenchLeader

GDP.pdf

Professional document tasks over PDFs.

As of 19 Sept 2026, GPT-5.6 Sol leads GDP.pdf on BenchLeader with 30.7%, ahead of GPT-5.6 Sol at 30.7%, across 39 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Knowledge
Index weight
Reference only
Models
39
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Economically valuable tasks that hinge on reading and producing PDF documents.

How it is scored

Score, as published.

What to keep in mind

New; small set.

39 of 39
#
1GPT-5.6 SolmaxOpenAI30.7%68.82026-07-09
2GPT-5.6 SolOpenAI30.7%56.82026-07-09
3Claude Fable 5Anthropic30.0%68.32026-06-09
4Claude Fable 5maxAnthropic29.8%65.82026-06-09
5GPT-5.5xhighOpenAI26.0%67.62026-04-23
6GPT-5.6 TerramediumOpenAI24.7%58.42026-07-09
7GPT-5.6 TerraOpenAI24.7%2026-07-09
8Claude Opus 5maxAnthropic24.0%69.92026-07-24
9Claude Opus 4.8maxAnthropic24.0%64.22026-05-28
10Gemini 3.7 FlashhighGoogle23.8%64.62026-08-13
11Qwen3 8xhighAlibaba23.2%57.72026-08-02
12GPT-5.6 LunamediumOpenAI22.7%54.02026-07-09
13GPT-5.6 LunaOpenAI22.7%2026-07-09
14Gemini 3.7 FlashmediumGoogle21.8%64.72026-08-13
15Claude Opus 4.7maxAnthropic21.0%64.32026-04-16
16Kimi K3maxMoonshot AIopen ↗19.0%67.22026-07-16
17Claude Sonnet 4.6maxAnthropic18.0%57.62026-02-17
18Grok 4.6xhighxAI17.2%64.62026-08-12
19Gemini 3.1 ProGoogle17.0%63.92026-02-19
20Gemini 3.1 ProhighGoogle17.0%60.72026-02-19
21Grok 4.6highxAI16.0%64.52026-08-12
22Muse Spark 1.2xhighMeta16.0%64.12026-08-05
23Muse Spark 1.1Meta15.0%65.22026-07-09
24Muse Spark 1.1mediumMeta15.0%2026-07-09
25Gemini 3.5 FlashmediumGoogle14.0%64.22026-05-19
26Gemini 3.6 FlashhighGoogle14.0%61.72026-07-21
27Grok 4.5highxAI14.0%61.02026-07-08
28GLM-5.3-FlashmaxZhipu AIopen ↗14.0%55.32026-08-20
29Gemini 3.6 FlashGoogle14.0%51.42026-07-21
30Gemini 3.5 FlashGoogle14.0%51.32026-05-19
31Kimi K2.6Moonshot AIopen ↗12.0%60.52026-04-20
32Muse Spark 1.2Meta12.0%2026-08-05
33Gemini 3 FlashhighGoogle10.0%58.72025-12-17
34Gemini 3 FlashGoogle10.0%57.72025-12-17
35Gemini 3.5 Flash LitehighGoogle10.0%44.32026-07-21
36Grok 4.3highxAI8.0%58.22026-04-17
37Grok 4.3xAI8.0%50.52026-04-17
38Nova 2.0 Pro Previewno reasoningAmazon2.0%44.42025-12-02
39Nova 2.0 Pro PreviewAmazon2.0%2025-12-02

Cite as: BenchLeader, “GDP.pdf leaderboard”, https://www.benchleader.com/benchmarks/gdp_pdf, data as of 19 Sept 2026.

GDP.pdf: questions

What does GDP.pdf measure?
Economically valuable tasks that hinge on reading and producing PDF documents. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads GDP.pdf?
GPT-5.6 Sol leads GDP.pdf with 30.7% as of 19 Sept 2026, ahead of GPT-5.6 Sol at 30.7%.
How many models have GDP.pdf results?
39 model configurations have a GDP.pdf result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs GDP.pdf and how often is it updated?
GDP.pdf is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does GDP.pdf count toward the BenchLeader Index?
No. GDP.pdf is shown for reference but left out of the composite index.