LMArena Longer Queries
Text arena rating on long prompts.
As of 19 Sept 2026, Claude Opus 4.6 leads LMArena Longer Queries on BenchLeader with 1524, ahead of Claude Fable 5 at 1521, across 344 model configurations with a published result.
- Published by
- LMArena
- Category
- Human preference
- Index weight
- Reference only
- Models
- 344
- Data as of
- 19 Sept 2026
CC BY 4.0 — lmarena-ai/leaderboard-dataset on Hugging Face.
What the test looks like
Pairwise votes on prompts above a length threshold.
How it is scored
Bradley-Terry rating from pairwise votes, in Elo-like units, with style control.
What to keep in mind
Fewer votes than the overall board.
- 1Claude Opus 4.6 (high)1524
- 2Claude Fable 51521
- 3Claude Opus 4.61517
- 4Claude Opus 4.7 (high)1514
- 5Claude Fable 5.1 (max)1511
- 6Gemini 3.8 Flash (high)1509
- 7Claude Opus 5 (high)1508
- 8Claude Opus 4.71507
- 9Muse Spark 1.3 (max)1506
- 10Muse Spark 1.2 (xhigh)1502
- 11Claude Opus 4.8 (high)1501
- 12Kimi K3 (max)1500
- 13Gemini 3.1 Pro1500
- 14Claude Opus 4.81500
- 15Claude Opus 5 (max)1500
344 of 344
| # | ||||
|---|---|---|---|---|
| 1 | 1524 | 61.3 | 2026-09-13 | |
| 2 | 1521 | 68.3 | 2026-09-13 | |
| 3 | 1517 | 63.7 | 2026-09-13 | |
| 4 | 1514 | 62.4 | 2026-09-13 | |
| 5 | 1511 | 69.7 | 2026-09-13 | |
| 6 | 1509 | 64.5 | 2026-09-13 | |
| 7 | 1508 | 70.2 | 2026-09-13 | |
| 8 | 1507 | 64.5 | 2026-09-13 | |
| 9 | 1506 | 69.3 | 2026-09-13 | |
| 10 | 1502 | 64.1 | 2026-09-13 | |
| 11 | 1501 | 62.2 | 2026-09-13 | |
| 12 | 1500 | 67.2 | 2026-09-13 | |
| 13 | 1500 | 63.9 | 2026-09-13 | |
| 14 | 1500 | 61.9 | 2026-09-13 | |
| 15 | 1500 | 69.9 | 2026-09-13 | |
| 16 | 1498 | 64.6 | 2026-09-13 | |
| 17 | 1498 | 68.2 | 2026-09-13 | |
| 18 | 1494 | 57.7 | 2026-09-13 | |
| 19 | 1494 | 59.0 | 2026-09-13 | |
| 20 | 1494 | 61.0 | 2026-09-13 | |
| 21 | 1493 | 58.2 | 2026-09-13 | |
| 22 | 1492 | 65.8 | 2026-09-13 | |
| 23 | 1492 | 66.5 | 2026-09-13 | |
| 24 | 1491 | 61.1 | 2026-09-13 | |
| 25 | 1488 | 67.0 | 2026-09-13 | |
| 26 | 1487 | 61.7 | 2026-09-13 | |
| 27 | 1487 | 59.2 | 2026-09-13 | |
| 28 | 1486 | 58.8 | 2026-09-13 | |
| 29 | 1485 | 63.2 | 2026-09-13 | |
| 30 | 1485 | 57.2 | 2026-09-13 | |
| 31 | 1483 | 63.9 | 2026-09-13 | |
| 32 | 1483 | 59.0 | 2026-09-13 | |
| 33 | 1483 | 60.3 | 2026-09-13 | |
| 34 | 1483 | 54.3 | 2026-09-13 | |
| 35 | 1482 | 62.2 | 2026-09-13 | |
| 36 | 1482 | – | 2026-09-13 | |
| 37 | 1481 | 71.8 | 2026-09-13 | |
| 38 | 1481 | 63.6 | 2026-09-13 | |
| 39 | 1481 | 65.2 | 2026-09-13 | |
| 40 | 1480 | 56.9 | 2026-09-13 | |
| 41 | 1478 | 59.2 | 2026-09-13 | |
| 42 | 1478 | 64.2 | 2026-09-13 | |
| 43 | 1477 | 57.7 | 2026-09-13 | |
| 44 | 1476 | 63.9 | 2026-09-13 | |
| 45 | 1476 | 60.5 | 2026-09-13 | |
| 46 | 1475 | 59.2 | 2026-09-13 | |
| 47 | 1474 | 65.7 | 2026-09-13 | |
| 48 | 1474 | 63.3 | 2026-09-13 | |
| 49 | 1474 | 60.8 | 2026-09-13 | |
| 50 | 1474 | 58.4 | 2026-09-13 | |
| 51 | 1472 | 51.9 | 2026-09-13 | |
| 52 | 1472 | 52.6 | 2026-09-13 | |
| 53 | 1471 | 64.5 | 2026-09-13 | |
| 54 | 1470 | 54.7 | 2026-09-13 | |
| 55 | 1468 | 64.3 | 2026-09-13 | |
| 56 | 1468 | 52.7 | 2026-09-13 | |
| 57 | 1467 | 59.8 | 2026-09-13 | |
| 58 | 1466 | – | 2026-09-13 | |
| 59 | 1466 | 55.7 | 2026-09-13 | |
| 60 | 1465 | 59.9 | 2026-09-13 |
Cite as: BenchLeader, “LMArena Longer Queries leaderboard”, https://www.benchleader.com/benchmarks/lmarena_text_longer, data as of 19 Sept 2026.
LMArena Longer Queries: questions
- What does LMArena Longer Queries measure?
- Pairwise votes on prompts above a length threshold. Scores are reported in Elo-style ratings from pairwise comparisons; higher is better.
- Which AI model leads LMArena Longer Queries?
- Claude Opus 4.6 leads LMArena Longer Queries with 1524 as of 19 Sept 2026, ahead of Claude Fable 5 at 1521.
- How many models have LMArena Longer Queries results?
- 344 model configurations have a LMArena Longer Queries result on BenchLeader, all taken from LMArena.
- Who runs LMArena Longer Queries and how often is it updated?
- LMArena Longer Queries is published by LMArena. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does LMArena Longer Queries count toward the BenchLeader Index?
- No. LMArena Longer Queries is shown for reference but left out of the composite index.