BenchLeader

LMArena Longer Queries

Text arena rating on long prompts.

As of 19 Sept 2026, Claude Opus 4.6 leads LMArena Longer Queries on BenchLeader with 1524, ahead of Claude Fable 5 at 1521, across 344 model configurations with a published result.

Published by
LMArena
Category
Human preference
Index weight
Reference only
Models
344
Data as of
19 Sept 2026

CC BY 4.0 — lmarena-ai/leaderboard-dataset on Hugging Face.

What the test looks like

Pairwise votes on prompts above a length threshold.

How it is scored

Bradley-Terry rating from pairwise votes, in Elo-like units, with style control.

What to keep in mind

Fewer votes than the overall board.

344 of 344
#
1Claude Opus 4.6highAnthropic152461.32026-09-13
2Claude Fable 5Anthropic152168.32026-09-13
3Claude Opus 4.6Anthropic151763.72026-09-13
4Claude Opus 4.7highAnthropic151462.42026-09-13
5Claude Fable 5.1maxAnthropic151169.72026-09-13
6Gemini 3.8 FlashhighGoogle150964.52026-09-13
7Claude Opus 5highAnthropic150870.22026-09-13
8Claude Opus 4.7Anthropic150764.52026-09-13
9Muse Spark 1.3maxMeta150669.32026-09-13
10Muse Spark 1.2xhighMeta150264.12026-09-13
11Claude Opus 4.8highAnthropic150162.22026-09-13
12Kimi K3maxMoonshot AIopen ↗150067.22026-09-13
13Gemini 3.1 ProGoogle150063.92026-09-13
14Claude Opus 4.8Anthropic150061.92026-09-13
15Claude Opus 5maxAnthropic150069.92026-09-13
16Gemini 3.7 FlashhighGoogle149864.62026-09-13
17GPT-5.6 SolxhighOpenAI149868.22026-09-13
18Claude Opus 4.5highAnthropic149457.72026-09-13
19Claude Sonnet 4.6Anthropic149459.02026-09-13
20Qwen3 7maxAlibaba149461.02026-09-13
21Claude Opus 4.5Anthropic149358.22026-09-13
22GLM-5.3maxZhipu AIopen ↗149265.82026-09-13
23Qwen3 8maxAlibaba149266.52026-09-13
24Gemini 3 ProGoogle149161.12026-09-13
25GPT-5.5highOpenAI148867.02026-09-13
26Gemini 3.6 FlashhighGoogle148761.72026-09-13
27MiMo-V2.5-ProXiaomiopen ↗148759.22026-09-13
28Claude Sonnet 4.5highAnthropic148658.82026-09-13
29GPT-5.5OpenAI148563.22026-09-13
30Claude Opus 4.1thinkingAnthropic148557.22026-09-13
31GLM-5.2maxZhipu AIopen ↗148363.92026-09-13
32GPT-5.4highOpenAI148359.02026-09-13
33Grok 4.5xAI148360.32026-09-13
34Claude Sonnet 4.5Anthropic148354.32026-09-13
35Claude Sonnet 5highAnthropic148262.22026-09-13
36Qwen3 5maxAlibaba14822026-09-13
37GPT-6 AstramaxOpenAI148171.82026-09-13
38Gemini 3.5 FlashhighGoogle148163.62026-09-13
39Muse Spark 1.1Meta148165.22026-09-13
40GLM-5.1Zhipu AIopen ↗148056.92026-09-13
41GPT-5.4OpenAI147859.22026-09-13
42Gemini 3.5 FlashmediumGoogle147864.22026-09-13
43Gemini 3 FlashGoogle147757.72026-09-13
44GLM-5.3-FlashZhipu AIopen ↗147663.92026-09-13
45Kimi K2.6Moonshot AIopen ↗147660.52026-09-13
46GPT-5.5 InstantOpenAI147559.22026-09-13
47Muse SparkMeta147465.72026-09-13
48Qwen3 6maxAlibaba147463.32026-09-13
49Deepseek v4 ProhighDeepSeek147460.82026-09-13
50GPT-5.2OpenAI147458.42026-09-13
51Claude Opus 4.1Anthropic147251.92026-09-13
52DeepSeek V4 ProDeepSeekopen ↗147252.62026-09-13
53Grok 4.6highxAI147164.52026-09-13
54GLM-5Zhipu AIopen147054.72026-09-13
55GPT-5.6 TerraxhighOpenAI146864.32026-09-13
56Gemma 4 31BGoogleopen ↗146852.72026-09-13
57Mimo v2 ProXiaomi146759.82026-09-13
58Ernie 5.1Baidu14662026-09-13
59Claude Opus 4thinkingAnthropic146655.72026-09-13
60Grok 4.20thinkingxAI146559.92026-09-13

Cite as: BenchLeader, “LMArena Longer Queries leaderboard”, https://www.benchleader.com/benchmarks/lmarena_text_longer, data as of 19 Sept 2026.

LMArena Longer Queries: questions

What does LMArena Longer Queries measure?
Pairwise votes on prompts above a length threshold. Scores are reported in Elo-style ratings from pairwise comparisons; higher is better.
Which AI model leads LMArena Longer Queries?
Claude Opus 4.6 leads LMArena Longer Queries with 1524 as of 19 Sept 2026, ahead of Claude Fable 5 at 1521.
How many models have LMArena Longer Queries results?
344 model configurations have a LMArena Longer Queries result on BenchLeader, all taken from LMArena.
Who runs LMArena Longer Queries and how often is it updated?
LMArena Longer Queries is published by LMArena. BenchLeader re-reads the published results every morning and records the date each result was published.
Does LMArena Longer Queries count toward the BenchLeader Index?
No. LMArena Longer Queries is shown for reference but left out of the composite index.