BenchLeader

Harvey's Legal Agent Benchmark

Agentic legal work over files, written with Harvey. Run by Vals AI.

As of 19 Sept 2026, Muse Spark 1.2 leads Harvey's Legal Agent Benchmark on BenchLeader with 25.4%, ahead of Muse Spark 1.3 at 23.8%, across 59 model configurations with a published result.

Published by
Vals AI
Category
Agents & tools
Index weight
Reference only
Models
59
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

Legal tasks that require working across a set of documents, as an associate would, from drafting to review.

How it is scored

Rubric score, run by Vals AI with Harvey.

What to keep in mind

Co-designed with a legal-AI vendor, which Vals discloses.

59 of 59
#
1Muse Spark 1.2xhighMeta25.4%64.12026-09-15
2Muse Spark 1.3maxMeta23.8%69.32026-09-15
3Muse Spark 1.3xhighMeta22.1%68.02026-09-15
4Muse Spark 1.1xhighMeta20.0%61.82026-09-15
5Grok 4.6highxAI15.8%64.52026-09-15
6Grok 4.5highxAI12.9%61.02026-09-15
7Claude Fable 5Anthropic11.3%68.32026-09-15
8Qwen3.8 27BxhighAlibabaopen ↗11.3%59.22026-09-15
9Kimi K3Moonshot AIopen ↗10.8%64.22026-09-15
10Qwen3 8maxAlibaba10.4%66.52026-09-15
11Gemini 3.8 FlashhighGoogle10.0%64.52026-09-15
12Claude Opus 4.8Anthropic9.6%61.92026-09-15
13Gemini 3.7 FlashhighGoogle8.8%64.62026-09-15
14GLM-5.3maxZhipu AIopen ↗8.3%65.82026-09-15
15DeepSeek V4 Flash 0731highDeepSeekopen ↗8.3%55.02026-09-15
16DeepSeek V4 Pro 0813maxDeepSeekopen ↗7.5%56.02026-09-15
17GLM-5.2maxZhipu AIopen ↗7.1%63.92026-09-15
18Claude Opus 5Anthropic6.7%66.72026-09-15
19Claude Opus 4.7Anthropic6.7%64.52026-09-15
20Claude Fable 5.1Anthropic6.7%64.52026-09-15
21GLM-5.3-FlashmaxZhipu AIopen ↗6.7%55.32026-09-15
22DeepSeek V4.1 FlashhighDeepSeekopen ↗6.7%2026-09-15
23GPT-6 AstramaxOpenAI5.4%71.82026-09-15
24Claude Sonnet 4.6Anthropic5.0%59.02026-09-15
25Claude Sonnet 5Anthropic5.0%57.22026-09-15
26MiniMax-M3MiniMaxopen ↗4.2%57.52026-09-15
27GPT-5.5xhighOpenAI3.8%67.62026-09-15
28DeepSeek V4 PromaxDeepSeekopen ↗3.8%64.22026-09-15
29Gemini 3.6 FlashhighGoogle3.3%61.72026-09-15
30GPT-5.6 SolmaxOpenAI2.5%68.82026-09-15
31Gemini 3.5 FlashhighGoogle2.5%63.62026-09-15
32MiMo-V2.5-ProXiaomiopen ↗2.1%59.22026-09-15
33Qwen3 7maxAlibaba1.7%61.02026-09-15
34Kimi K2.6Moonshot AIopen ↗1.7%60.52026-09-15
35Inkling SmallThinking Machinesopen ↗1.7%56.32026-09-15
36GPT-5.6 LunamaxOpenAI1.3%60.22026-09-15
37Qwen3.6 PlusAlibabaopen1.3%58.52026-09-15
38Ling 3.0 FlashAnt Groupopen ↗1.3%53.62026-09-15
39GPT-5.6 TerramaxOpenAI0.8%65.02026-09-15
40Claude Haiku 4.5thinkingAnthropic0.8%47.52026-09-15
41Grok 4.3highxAI0.4%58.22026-09-15
42Nemotron 3 Ultra 550B A55BNVIDIAopen ↗0.4%48.12026-09-15
43Mistral Medium 3.5highMistral AIopen0.4%2026-09-15
44GPT-5.4xhighOpenAI0.0%65.42026-09-15
45Gemini 3.1 ProhighGoogle0.0%60.72026-09-15
46Grok 4.20thinkingxAI0.0%59.92026-09-15
47Qwen3.7 PlusAlibaba0.0%59.92026-09-15
48Gemini 3 FlashhighGoogle0.0%58.72026-09-15
49Kimi K2.5thinkingMoonshot AIopen0.0%58.52026-09-15
50GLM-5.1Zhipu AIopen ↗0.0%56.92026-09-15
51MiniMax-M2.7MiniMaxopen ↗0.0%56.12026-09-15
52GPT-5.4 minixhighOpenAI0.0%55.72026-09-15
53GPT-5.4 nanohighOpenAI0.0%49.12026-09-15
54Gemini 3.1 Flash LitehighGoogle0.0%49.12026-09-15
55Gemini 3.5 Flash LitehighGoogle0.0%44.32026-09-15
56Laguna M.1Poolsideopen0.0%37.72026-09-15
57Laguna XS.2Poolside0.0%36.52026-09-15
58Mercury 2.5highInception0.0%2026-09-15
59Nemotron Lightning 3.5 30B A3bNVIDIA0.0%2026-09-15

Cite as: BenchLeader, “Harvey's Legal Agent Benchmark leaderboard”, https://www.benchleader.com/benchmarks/vals_hlab, data as of 19 Sept 2026.

Harvey's Legal Agent Benchmark: questions

What does Harvey's Legal Agent Benchmark measure?
Legal tasks that require working across a set of documents, as an associate would, from drafting to review. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads Harvey's Legal Agent Benchmark?
Muse Spark 1.2 leads Harvey's Legal Agent Benchmark with 25.4% as of 19 Sept 2026, ahead of Muse Spark 1.3 at 23.8%.
How many models have Harvey's Legal Agent Benchmark results?
59 model configurations have a Harvey's Legal Agent Benchmark result on BenchLeader, all taken from Vals AI.
Who runs Harvey's Legal Agent Benchmark and how often is it updated?
Harvey's Legal Agent Benchmark is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does Harvey's Legal Agent Benchmark count toward the BenchLeader Index?
No. Harvey's Legal Agent Benchmark is shown for reference but left out of the composite index.