LMArena Search
Pairwise votes on answers that used web search.
As of 19 Sept 2026, GPT-5.6 Sol leads LMArena Search on BenchLeader with 1256, ahead of Claude Fable 5 at 1229, across 33 model configurations with a published result.
- Published by
- LMArena
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 33
- Data as of
- 19 Sept 2026
CC BY 4.0 — lmarena-ai/leaderboard-dataset on Hugging Face.
What the test looks like
Search Arena compares models that answer with live web search, judged on the answer and its sourcing.
How it is scored
Bradley-Terry rating from pairwise votes, in Elo-like units, with style control.
What to keep in mind
Only search-enabled products are listed, and results depend on the day.
- 1GPT-5.6 Sol (xhigh)1256
- 2Claude Fable 51229
- 3GPT 5.5 Search1224
- 4Claude Opus 4.6 Search1223
- 5Claude Opus 4.71212
- 6Gemini 3.1 Pro Grounding1208
- 7Grok 4.20 Multi-Agent1205
- 8Grok 4.51202
- 9Gemini 3 Pro Grounding1201
- 10Claude Sonnet 4.6 Search1201
- 11Grok 4.20 Beta11197
- 12Claude Sonnet 5 Search1196
- 13Claude Opus 4.81195
- 14GPT 5.4 Search1195
- 15Grok 4.1 Fast Search1194
33 of 33
| # | ||||
|---|---|---|---|---|
| 1 | 1256 | 68.2 | 2026-08-24 | |
| 2 | 1229 | 68.3 | 2026-08-24 | |
| 3 | 1224 | – | 2026-08-24 | |
| 4 | 1223 | – | 2026-08-24 | |
| 5 | 1212 | 64.5 | 2026-08-24 | |
| 6 | 1208 | – | 2026-08-24 | |
| 7 | 1205 | 59.2 | 2026-08-24 | |
| 8 | 1202 | 60.3 | 2026-08-24 | |
| 9 | 1201 | – | 2026-08-24 | |
| 10 | 1201 | – | 2026-08-24 | |
| 11 | 1197 | – | 2026-08-24 | |
| 12 | 1196 | – | 2026-08-24 | |
| 13 | 1195 | 61.9 | 2026-08-24 | |
| 14 | 1195 | – | 2026-08-24 | |
| 15 | 1194 | – | 2026-08-24 | |
| 16 | 1193 | – | 2026-08-24 | |
| 17 | 1193 | 50.5 | 2026-08-24 | |
| 18 | 1192 | – | 2026-08-24 | |
| 19 | 1191 | – | 2026-08-24 | |
| 20 | 1187 | – | 2026-08-24 | |
| 21 | 1183 | – | 2026-08-24 | |
| 22 | 1183 | – | 2026-08-24 | |
| 23 | 1175 | – | 2026-08-24 | |
| 24 | 1167 | – | 2026-08-24 | |
| 25 | 1162 | – | 2026-08-24 | |
| 26 | 1160 | – | 2026-08-24 | |
| 27 | 1151 | – | 2026-08-24 | |
| 28 | 1143 | – | 2026-08-24 | |
| 29 | 1142 | – | 2026-08-24 | |
| 30 | 1137 | – | 2026-08-24 | |
| 31 | 1122 | – | 2026-08-24 | |
| 32 | 1081 | – | 2026-08-24 | |
| 33 | Diffbot Small XlDiffbot | 1062 | – | 2026-08-24 |
Cite as: BenchLeader, “LMArena Search leaderboard”, https://www.benchleader.com/benchmarks/lmarena_search, data as of 19 Sept 2026.
LMArena Search: questions
- What does LMArena Search measure?
- Search Arena compares models that answer with live web search, judged on the answer and its sourcing. Scores are reported in Elo-style ratings from pairwise comparisons; higher is better.
- Which AI model leads LMArena Search?
- GPT-5.6 Sol leads LMArena Search with 1256 as of 19 Sept 2026, ahead of Claude Fable 5 at 1229.
- How many models have LMArena Search results?
- 33 model configurations have a LMArena Search result on BenchLeader, all taken from LMArena.
- Who runs LMArena Search and how often is it updated?
- LMArena Search is published by LMArena. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does LMArena Search count toward the BenchLeader Index?
- No. LMArena Search is shown for reference but left out of the composite index.