BenchLeader

LMArena Instruction Following

Text arena rating on prompts with explicit instructions.

As of 19 Sept 2026, Claude Opus 4.6 leads LMArena Instruction Following on BenchLeader with 1514, ahead of Claude Fable 5 at 1511, across 366 model configurations with a published result.

Published by
LMArena
Category
Instruction following
Index weight
Reference only
Models
366
Data as of
19 Sept 2026

CC BY 4.0 — lmarena-ai/leaderboard-dataset on Hugging Face.

What the test looks like

Pairwise votes on prompts that set explicit constraints on the answer.

How it is scored

Bradley-Terry rating from pairwise votes, in Elo-like units, with style control.

What to keep in mind

Voters judge compliance by eye; see IFEval for programmatic checks.

366 of 366
#
1Claude Opus 4.6highAnthropic151461.32026-09-13
2Claude Fable 5Anthropic151168.32026-09-13
3Claude Opus 4.7highAnthropic150362.42026-09-13
4Claude Opus 4.6Anthropic150063.72026-09-13
5Claude Opus 5highAnthropic149970.22026-09-13
6Claude Fable 5.1maxAnthropic149669.72026-09-13
7Claude Opus 5maxAnthropic149369.92026-09-13
8Claude Opus 4.7Anthropic149364.52026-09-13
9Gemini 3.8 FlashhighGoogle149264.52026-09-13
10Claude Opus 4.8highAnthropic149162.22026-09-13
11GPT-5.6 SolxhighOpenAI148768.22026-09-13
12Gemini 3.7 FlashhighGoogle148464.62026-09-13
13Kimi K3maxMoonshot AIopen ↗148467.22026-09-13
14Claude Opus 4.5highAnthropic148357.72026-09-13
15Muse Spark 1.3maxMeta148369.32026-09-13
16Gemini 3.1 ProGoogle148163.92026-09-13
17GLM-5.3maxZhipu AIopen ↗148065.82026-09-13
18Muse Spark 1.2xhighMeta148064.12026-09-13
19GPT-5.5highOpenAI147867.02026-09-13
20Qwen3 8maxAlibaba147766.52026-09-13
21Claude Opus 4.8Anthropic147761.92026-09-13
22Claude Sonnet 4.6Anthropic147659.02026-09-13
23Claude Opus 4.5Anthropic147558.22026-09-13
24GPT-5.5OpenAI147363.22026-09-13
25Gemini 3.6 FlashhighGoogle147361.72026-09-13
26Gemini 3 ProGoogle147361.12026-09-13
27GPT-5.4highOpenAI147259.02026-09-13
28Muse Spark 1.1Meta147165.22026-09-13
29GLM-5.3-FlashZhipu AIopen ↗147163.92026-09-13
30MiMo-V2.5-ProXiaomiopen ↗147059.22026-09-13
31Qwen3 5maxAlibaba14692026-09-13
32GLM-5.2maxZhipu AIopen ↗146663.92026-09-13
33Grok 4.5xAI146660.32026-09-13
34Claude Sonnet 5highAnthropic146562.22026-09-13
35Gemini 3.5 FlashhighGoogle146563.62026-09-13
36Qwen3 7maxAlibaba146461.02026-09-13
37Claude Sonnet 4.5highAnthropic146458.82026-09-13
38Muse SparkMeta146365.72026-09-13
39Gemini 3.5 FlashmediumGoogle146364.22026-09-13
40Claude Sonnet 4.5Anthropic146254.32026-09-13
41GPT-5.6 TerraxhighOpenAI146264.32026-09-13
42GPT-5.4OpenAI146159.22026-09-13
43GPT-6 AstramaxOpenAI146171.82026-09-13
44GLM-5.1Zhipu AIopen ↗146156.92026-09-13
45Claude Opus 4.1thinkingAnthropic145957.22026-09-13
46Gemini 3 FlashGoogle145957.72026-09-13
47GPT-5.2OpenAI145758.42026-09-13
48Deepseek v4 ProhighDeepSeek145760.82026-09-13
49GPT-5.5 InstantOpenAI145659.22026-09-13
50Claude Opus 4.1Anthropic145551.92026-09-13
51Kimi K2.6Moonshot AIopen ↗145460.52026-09-13
52Ernie 5.1Baidu14532026-09-13
53Grok 4.6highxAI145364.52026-09-13
54DeepSeek V4 ProDeepSeekopen ↗145352.62026-09-13
55Gemma 4 31BGoogleopen ↗145252.72026-09-13
56GPT-5.1highOpenAI145058.72026-09-13
57Grok 4.20 Beta1xAI14502026-09-13
58Qwen3 6maxAlibaba144963.32026-09-13
59Qwen3.7 PlusAlibaba144759.92026-09-13
60GLM-5Zhipu AIopen144654.72026-09-13

Cite as: BenchLeader, “LMArena Instruction Following leaderboard”, https://www.benchleader.com/benchmarks/lmarena_text_if, data as of 19 Sept 2026.

LMArena Instruction Following: questions

What does LMArena Instruction Following measure?
Pairwise votes on prompts that set explicit constraints on the answer. Scores are reported in Elo-style ratings from pairwise comparisons; higher is better.
Which AI model leads LMArena Instruction Following?
Claude Opus 4.6 leads LMArena Instruction Following with 1514 as of 19 Sept 2026, ahead of Claude Fable 5 at 1511.
How many models have LMArena Instruction Following results?
366 model configurations have a LMArena Instruction Following result on BenchLeader, all taken from LMArena.
Who runs LMArena Instruction Following and how often is it updated?
LMArena Instruction Following is published by LMArena. BenchLeader re-reads the published results every morning and records the date each result was published.
Does LMArena Instruction Following count toward the BenchLeader Index?
No. LMArena Instruction Following is shown for reference but left out of the composite index.