BenchLeader

LMArena Document

Pairwise votes on answers about uploaded documents.

As of 19 Sept 2026, Claude Fable 5 leads LMArena Document on BenchLeader with 1499, ahead of Claude Opus 4.6 at 1495, across 44 model configurations with a published result.

Published by
LMArena
Category
Long context
Index weight
Reference only
Models
44
Data as of
19 Sept 2026

CC BY 4.0 — lmarena-ai/leaderboard-dataset on Hugging Face.

What the test looks like

Document Arena compares models answering questions over an uploaded file.

How it is scored

Bradley-Terry rating from pairwise votes, in Elo-like units, with style control.

What to keep in mind

Fewer votes than the text arena; document types are whatever users upload.

44 of 44
#
1Claude Fable 5Anthropic149968.32026-09-13
2Claude Opus 4.6highAnthropic149561.32026-09-13
3Claude Opus 4.7Anthropic149464.52026-09-13
4Claude Opus 4.6Anthropic149463.72026-09-13
5Claude Opus 4.7highAnthropic149362.42026-09-13
6Claude Fable 5.1maxAnthropic149369.72026-09-13
7Claude Opus 5highAnthropic149070.22026-09-13
8GPT-5.6 SolxhighOpenAI148968.22026-09-13
9Claude Opus 4.8highAnthropic148662.22026-09-13
10GPT-5.5OpenAI148363.22026-09-13
11GPT-5.5highOpenAI148067.02026-09-13
12Muse Spark 1.1Meta148065.22026-09-13
13Claude Sonnet 4.6Anthropic147859.02026-09-13
14GPT-5.6 TerraxhighOpenAI147764.32026-09-13
15Claude Sonnet 5highAnthropic147662.22026-09-13
16Claude Opus 4.8Anthropic147561.92026-09-13
17GPT-6 AstramaxOpenAI147371.82026-09-13
18GPT-5.4OpenAI147059.22026-09-13
19Muse Spark 1.3maxMeta146869.32026-09-13
20Muse SparkMeta146765.72026-09-13
21Claude Opus 4.5Anthropic146558.22026-09-13
22Grok 4.5xAI146460.32026-09-13
23GPT-5.6 LunaxhighOpenAI146261.12026-09-13
24Grok 4.6highxAI146164.52026-09-13
25Gemini 3.5 FlashmediumGoogle146164.22026-09-13
26GPT-5.5 InstantOpenAI146059.22026-09-13
27Gemini 3.1 ProGoogle145963.92026-09-13
28Gemini 3.5 FlashhighGoogle145363.62026-09-13
29Gemini 3.6 FlashhighGoogle145361.72026-09-13
30Qwen3.7 PlusAlibaba145359.92026-09-13
31Gemini 3 ProGoogle145161.12026-09-13
32Kimi K2.6Moonshot AIopen ↗145060.52026-09-13
33Claude Sonnet 4.5Anthropic144954.32026-09-13
34Gemma 4 31BGoogleopen ↗144552.72026-09-13
35Grok 4.20thinkingxAI143959.92026-09-13
36Gemini 3 FlashGoogle143957.72026-09-13
37Kimi K2.5thinkingMoonshot AIopen143658.52026-09-13
38MiniMax-M3MiniMaxopen ↗143357.52026-09-13
39Gemini 2.5 ProGoogle143154.52026-09-13
40GPT-5.2OpenAI142558.42026-09-13
41GPT-5.2highOpenAI142355.72026-09-13
42Claude Haiku 4.5Anthropic141948.12026-09-13
43GPT-5.1OpenAI141657.22026-09-13
44GLM-5V-TurboZhipu AI140056.72026-09-13

Cite as: BenchLeader, “LMArena Document leaderboard”, https://www.benchleader.com/benchmarks/lmarena_document, data as of 19 Sept 2026.

LMArena Document: questions

What does LMArena Document measure?
Document Arena compares models answering questions over an uploaded file. Scores are reported in Elo-style ratings from pairwise comparisons; higher is better.
Which AI model leads LMArena Document?
Claude Fable 5 leads LMArena Document with 1499 as of 19 Sept 2026, ahead of Claude Opus 4.6 at 1495.
How many models have LMArena Document results?
44 model configurations have a LMArena Document result on BenchLeader, all taken from LMArena.
Who runs LMArena Document and how often is it updated?
LMArena Document is published by LMArena. BenchLeader re-reads the published results every morning and records the date each result was published.
Does LMArena Document count toward the BenchLeader Index?
No. LMArena Document is shown for reference but left out of the composite index.