BenchLeader

LMArena Multi-turn

Text arena rating on conversations with more than one turn.

As of 19 Sept 2026, Muse Spark 1.2 leads LMArena Multi-turn on BenchLeader with 1520, ahead of Claude Fable 5 at 1518, across 364 model configurations with a published result.

Published by
LMArena
Category
Human preference
Index weight
Reference only
Models
364
Data as of
19 Sept 2026

CC BY 4.0 — lmarena-ai/leaderboard-dataset on Hugging Face.

What the test looks like

Pairwise votes on multi-turn conversations, rewarding models that track context.

How it is scored

Bradley-Terry rating from pairwise votes, in Elo-like units, with style control.

What to keep in mind

Fewer votes than the overall board.

364 of 364
#
1Muse Spark 1.2xhighMeta152064.12026-09-13
2Claude Fable 5Anthropic151868.32026-09-13
3Claude Opus 4.6highAnthropic151761.32026-09-13
4Claude Opus 4.7highAnthropic151662.42026-09-13
5Claude Opus 4.7Anthropic151564.52026-09-13
6Gemini 3.8 FlashhighGoogle151364.52026-09-13
7Claude Opus 4.6Anthropic151163.72026-09-13
8GPT-6 AstramaxOpenAI149971.82026-09-13
9Gemini 3.7 FlashhighGoogle149864.62026-09-13
10Muse Spark 1.1Meta149765.22026-09-13
11Claude Opus 4.8highAnthropic149662.22026-09-13
12Claude Opus 4.8Anthropic149661.92026-09-13
13Kimi K3maxMoonshot AIopen ↗149667.22026-09-13
14Gemini 3 ProGoogle149561.12026-09-13
15Gemini 3.1 ProGoogle149563.92026-09-13
16Qwen3 8maxAlibaba149566.52026-09-13
17GPT-5.4highOpenAI149459.02026-09-13
18GPT-5.2OpenAI149458.42026-09-13
19Muse Spark 1.3maxMeta149369.32026-09-13
20GLM-5.3maxZhipu AIopen ↗149365.82026-09-13
21Muse SparkMeta149165.72026-09-13
22GPT-5.5highOpenAI148967.02026-09-13
23Claude Fable 5.1maxAnthropic148869.72026-09-13
24GPT-5.6 SolxhighOpenAI148768.22026-09-13
25Claude Opus 4.5highAnthropic148657.72026-09-13
26Gemini 3.6 FlashhighGoogle148561.72026-09-13
27Claude Opus 5maxAnthropic148569.92026-09-13
28Claude Opus 4.5Anthropic148458.22026-09-13
29Claude Opus 5highAnthropic148470.22026-09-13
30Gemini 3.5 FlashhighGoogle148363.62026-09-13
31GPT-5.5 InstantOpenAI148359.22026-09-13
32Gemini 3 FlashGoogle148357.72026-09-13
33Grok 4.20 Beta1xAI14822026-09-13
34Qwen3 7maxAlibaba148261.02026-09-13
35GPT-5.5OpenAI148163.22026-09-13
36Claude Sonnet 4.6Anthropic148159.02026-09-13
37GPT-5.4OpenAI148159.22026-09-13
38Grok 4.20thinkingxAI148059.92026-09-13
39Qwen3 5maxAlibaba14792026-09-13
40Claude Sonnet 4.5Anthropic147854.32026-09-13
41MiMo-V2.5-ProXiaomiopen ↗147859.22026-09-13
42GLM-5.1Zhipu AIopen ↗147756.92026-09-13
43Grok 4.5xAI147660.32026-09-13
44Gemini 3.5 FlashmediumGoogle147664.22026-09-13
45DeepSeek V4 ProDeepSeekopen ↗147552.62026-09-13
46GLM-5.3-FlashZhipu AIopen ↗147563.92026-09-13
47Gemini 3 FlashminimalGoogle147454.92026-09-13
48Claude Sonnet 5highAnthropic147362.22026-09-13
49Grok 4.20 Multi-AgentxAI147359.22026-09-13
50GPT-5.6 TerraxhighOpenAI147364.32026-09-13
51GLM-5Zhipu AIopen147354.72026-09-13
52Claude Opus 4.1thinkingAnthropic147357.22026-09-13
53Ernie 5.1Baidu14732026-09-13
54GPT-4oOpenAI147144.32026-09-13
55Claude Sonnet 4.5highAnthropic147058.82026-09-13
56Deepseek v4 ProhighDeepSeek146960.82026-09-13
57GLM-5.2maxZhipu AIopen ↗146963.92026-09-13
58Qwen3 6maxAlibaba146963.32026-09-13
59Mimo v2 ProXiaomi146959.82026-09-13
60Claude Opus 4.1Anthropic146951.92026-09-13

Cite as: BenchLeader, “LMArena Multi-turn leaderboard”, https://www.benchleader.com/benchmarks/lmarena_text_multiturn, data as of 19 Sept 2026.

LMArena Multi-turn: questions

What does LMArena Multi-turn measure?
Pairwise votes on multi-turn conversations, rewarding models that track context. Scores are reported in Elo-style ratings from pairwise comparisons; higher is better.
Which AI model leads LMArena Multi-turn?
Muse Spark 1.2 leads LMArena Multi-turn with 1520 as of 19 Sept 2026, ahead of Claude Fable 5 at 1518.
How many models have LMArena Multi-turn results?
364 model configurations have a LMArena Multi-turn result on BenchLeader, all taken from LMArena.
Who runs LMArena Multi-turn and how often is it updated?
LMArena Multi-turn is published by LMArena. BenchLeader re-reads the published results every morning and records the date each result was published.
Does LMArena Multi-turn count toward the BenchLeader Index?
No. LMArena Multi-turn is shown for reference but left out of the composite index.