BenchLeader

Legal Research Bench

Agentic US legal research questions. Run by Vals AI.

As of 19 Sept 2026, Muse Spark 1.3 leads Legal Research Bench on BenchLeader with 55.3%, ahead of Claude Opus 5 at 55.3%, across 58 model configurations with a published result.

Published by
Vals AI
Category
Agents & tools
Index weight
Reference only
Models
58
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

The model must research a US legal question with tools and return a supported answer, the way an associate would.

How it is scored

Accuracy graded by lawyers, run by Vals AI.

What to keep in mind

US-specific and dependent on the research tools provided.

58 of 58
#
1Muse Spark 1.3maxMeta55.3%69.32026-09-15
2Claude Opus 5Anthropic55.3%66.72026-09-15
3Claude Fable 5.1Anthropic55.3%64.52026-09-15
4Claude Fable 5Anthropic49.5%68.32026-09-15
5GLM-5.3maxZhipu AIopen ↗49.0%65.82026-09-15
6GPT-5.6 SolmaxOpenAI48.1%68.82026-09-15
7Grok 4.6highxAI48.1%64.52026-09-15
8Qwen3 8maxAlibaba47.6%66.52026-09-15
9GLM-5.3-FlashmaxZhipu AIopen ↗45.2%55.32026-09-15
10Kimi K3Moonshot AIopen ↗44.2%64.22026-09-15
11Muse Spark 1.2xhighMeta43.8%64.12026-09-15
12Claude Opus 4.8Anthropic43.8%61.92026-09-15
13Claude Sonnet 5Anthropic41.8%57.22026-09-15
14GPT-5.6 TerramaxOpenAI41.4%65.02026-09-15
15DeepSeek V4.1 FlashhighDeepSeekopen ↗41.4%2026-09-15
16Muse Spark 1.3xhighMeta40.9%68.02026-09-15
17DeepSeek V4 Pro 0813maxDeepSeekopen ↗40.9%56.02026-09-15
18GPT-5.5xhighOpenAI40.4%67.62026-09-15
19GPT-6 AstramaxOpenAI39.4%71.82026-09-15
20Gemini 3.8 FlashhighGoogle38.9%64.52026-09-15
21Claude Opus 4.7Anthropic38.5%64.52026-09-15
22Claude Sonnet 4.6Anthropic38.5%59.02026-09-15
23Muse Spark 1.1xhighMeta38.0%61.82026-09-15
24Grok 4.5highxAI38.0%61.02026-09-15
25GPT-5.6 LunamaxOpenAI36.5%60.22026-09-15
26Qwen3.8 27BxhighAlibabaopen ↗36.1%59.22026-09-15
27Gemini 3.7 FlashhighGoogle34.6%64.62026-09-15
28GLM-5.2maxZhipu AIopen ↗31.3%63.92026-09-15
29Gemini 3.5 FlashhighGoogle30.8%63.62026-09-15
30DeepSeek V4 Flash 0731highDeepSeekopen ↗30.3%55.02026-09-15
31MiniMax-M3MiniMaxopen ↗29.8%57.52026-09-15
32GLM-5.1Zhipu AIopen ↗27.9%56.92026-09-15
33Qwen3 7maxAlibaba25.5%61.02026-09-15
34Inkling SmallThinking Machinesopen ↗25.5%56.32026-09-15
35Gemini 3.6 FlashhighGoogle25.0%61.72026-09-15
36DeepSeek V4 PromaxDeepSeekopen ↗23.1%64.22026-09-15
37Gemini 3.1 ProhighGoogle20.7%60.72026-09-15
38Gemini 3 FlashhighGoogle18.3%58.72026-09-15
39Qwen3.7 PlusAlibaba16.4%59.92026-09-15
40Kimi K2.6Moonshot AIopen ↗15.9%60.52026-09-15
41MiMo-V2.5-ProXiaomiopen ↗15.9%59.22026-09-15
42Kimi K2.5thinkingMoonshot AIopen15.9%58.52026-09-15
43Grok 4.3highxAI15.4%58.22026-09-15
44Nemotron 3 Ultra 550B A55BNVIDIAopen ↗15.4%48.12026-09-15
45Qwen3.6 PlusAlibabaopen14.9%58.52026-09-15
46Grok 4.20thinkingxAI13.9%59.92026-09-15
47Gemini 3.5 Flash LitehighGoogle13.9%44.32026-09-15
48GPT-5.4 minixhighOpenAI12.5%55.72026-09-15
49MiniMax-M2.7MiniMaxopen ↗10.6%56.12026-09-15
50Claude Haiku 4.5thinkingAnthropic10.6%47.52026-09-15
51Mistral Medium 3.5highMistral AIopen9.1%2026-09-15
52GPT-5.4 nanohighOpenAI6.3%49.12026-09-15
53Mercury 2.5highInception4.3%2026-09-15
54Gemini 3.1 Flash LiteGoogle3.4%54.52026-09-15
55Nemotron Lightning 3.5 30B A3bNVIDIA2.9%2026-09-15
56Laguna M.1Poolsideopen2.4%37.72026-09-15
57Laguna XS.2Poolside1.0%36.52026-09-15
58Ling 3.0 FlashAnt Groupopen ↗0.0%53.62026-09-15

Cite as: BenchLeader, “Legal Research Bench leaderboard”, https://www.benchleader.com/benchmarks/vals_legal_research, data as of 19 Sept 2026.

Legal Research Bench: questions

What does Legal Research Bench measure?
The model must research a US legal question with tools and return a supported answer, the way an associate would. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads Legal Research Bench?
Muse Spark 1.3 leads Legal Research Bench with 55.3% as of 19 Sept 2026, ahead of Claude Opus 5 at 55.3%.
How many models have Legal Research Bench results?
58 model configurations have a Legal Research Bench result on BenchLeader, all taken from Vals AI.
Who runs Legal Research Bench and how often is it updated?
Legal Research Bench is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does Legal Research Bench count toward the BenchLeader Index?
No. Legal Research Bench is shown for reference but left out of the composite index.