Harvey's Legal Agent Benchmark
Agentic legal work over files, written with Harvey. Run by Vals AI.
As of 19 Sept 2026, Muse Spark 1.2 leads Harvey's Legal Agent Benchmark on BenchLeader with 25.4%, ahead of Muse Spark 1.3 at 23.8%, across 59 model configurations with a published result.
- Published by
- Vals AI
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 59
- Data as of
- 19 Sept 2026
Vals AI (vals.ai).
What the test looks like
Legal tasks that require working across a set of documents, as an associate would, from drafting to review.
How it is scored
Rubric score, run by Vals AI with Harvey.
What to keep in mind
Co-designed with a legal-AI vendor, which Vals discloses.
- 1Muse Spark 1.2 (xhigh)25.4%
- 2Muse Spark 1.3 (max)23.8%
- 3Muse Spark 1.3 (xhigh)22.1%
- 4Muse Spark 1.1 (xhigh)20.0%
- 5Grok 4.6 (high)15.8%
- 6Grok 4.5 (high)12.9%
- 7Claude Fable 511.3%
- 8Qwen3.8 27B (xhigh)11.3%
- 9Kimi K310.8%
- 10Qwen3 8 (max)10.4%
- 11Gemini 3.8 Flash (high)10.0%
- 12Claude Opus 4.89.6%
- 13Gemini 3.7 Flash (high)8.8%
- 14GLM-5.3 (max)8.3%
- 15DeepSeek V4 Flash 0731 (high)8.3%
59 of 59
| # | ||||
|---|---|---|---|---|
| 1 | 25.4% | 64.1 | 2026-09-15 | |
| 2 | 23.8% | 69.3 | 2026-09-15 | |
| 3 | 22.1% | 68.0 | 2026-09-15 | |
| 4 | 20.0% | 61.8 | 2026-09-15 | |
| 5 | 15.8% | 64.5 | 2026-09-15 | |
| 6 | 12.9% | 61.0 | 2026-09-15 | |
| 7 | 11.3% | 68.3 | 2026-09-15 | |
| 8 | 11.3% | 59.2 | 2026-09-15 | |
| 9 | 10.8% | 64.2 | 2026-09-15 | |
| 10 | 10.4% | 66.5 | 2026-09-15 | |
| 11 | 10.0% | 64.5 | 2026-09-15 | |
| 12 | 9.6% | 61.9 | 2026-09-15 | |
| 13 | 8.8% | 64.6 | 2026-09-15 | |
| 14 | 8.3% | 65.8 | 2026-09-15 | |
| 15 | 8.3% | 55.0 | 2026-09-15 | |
| 16 | 7.5% | 56.0 | 2026-09-15 | |
| 17 | 7.1% | 63.9 | 2026-09-15 | |
| 18 | 6.7% | 66.7 | 2026-09-15 | |
| 19 | 6.7% | 64.5 | 2026-09-15 | |
| 20 | 6.7% | 64.5 | 2026-09-15 | |
| 21 | 6.7% | 55.3 | 2026-09-15 | |
| 22 | 6.7% | – | 2026-09-15 | |
| 23 | 5.4% | 71.8 | 2026-09-15 | |
| 24 | 5.0% | 59.0 | 2026-09-15 | |
| 25 | 5.0% | 57.2 | 2026-09-15 | |
| 26 | 4.2% | 57.5 | 2026-09-15 | |
| 27 | 3.8% | 67.6 | 2026-09-15 | |
| 28 | 3.8% | 64.2 | 2026-09-15 | |
| 29 | 3.3% | 61.7 | 2026-09-15 | |
| 30 | 2.5% | 68.8 | 2026-09-15 | |
| 31 | 2.5% | 63.6 | 2026-09-15 | |
| 32 | 2.1% | 59.2 | 2026-09-15 | |
| 33 | 1.7% | 61.0 | 2026-09-15 | |
| 34 | 1.7% | 60.5 | 2026-09-15 | |
| 35 | Inkling SmallThinking Machinesopen ↗ | 1.7% | 56.3 | 2026-09-15 |
| 36 | 1.3% | 60.2 | 2026-09-15 | |
| 37 | 1.3% | 58.5 | 2026-09-15 | |
| 38 | 1.3% | 53.6 | 2026-09-15 | |
| 39 | 0.8% | 65.0 | 2026-09-15 | |
| 40 | 0.8% | 47.5 | 2026-09-15 | |
| 41 | 0.4% | 58.2 | 2026-09-15 | |
| 42 | 0.4% | 48.1 | 2026-09-15 | |
| 43 | 0.4% | – | 2026-09-15 | |
| 44 | 0.0% | 65.4 | 2026-09-15 | |
| 45 | 0.0% | 60.7 | 2026-09-15 | |
| 46 | 0.0% | 59.9 | 2026-09-15 | |
| 47 | 0.0% | 59.9 | 2026-09-15 | |
| 48 | 0.0% | 58.7 | 2026-09-15 | |
| 49 | 0.0% | 58.5 | 2026-09-15 | |
| 50 | 0.0% | 56.9 | 2026-09-15 | |
| 51 | 0.0% | 56.1 | 2026-09-15 | |
| 52 | 0.0% | 55.7 | 2026-09-15 | |
| 53 | 0.0% | 49.1 | 2026-09-15 | |
| 54 | 0.0% | 49.1 | 2026-09-15 | |
| 55 | 0.0% | 44.3 | 2026-09-15 | |
| 56 | 0.0% | 37.7 | 2026-09-15 | |
| 57 | 0.0% | 36.5 | 2026-09-15 | |
| 58 | 0.0% | – | 2026-09-15 | |
| 59 | 0.0% | – | 2026-09-15 |
Cite as: BenchLeader, “Harvey's Legal Agent Benchmark leaderboard”, https://www.benchleader.com/benchmarks/vals_hlab, data as of 19 Sept 2026.
Harvey's Legal Agent Benchmark: questions
- What does Harvey's Legal Agent Benchmark measure?
- Legal tasks that require working across a set of documents, as an associate would, from drafting to review. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads Harvey's Legal Agent Benchmark?
- Muse Spark 1.2 leads Harvey's Legal Agent Benchmark with 25.4% as of 19 Sept 2026, ahead of Muse Spark 1.3 at 23.8%.
- How many models have Harvey's Legal Agent Benchmark results?
- 59 model configurations have a Harvey's Legal Agent Benchmark result on BenchLeader, all taken from Vals AI.
- Who runs Harvey's Legal Agent Benchmark and how often is it updated?
- Harvey's Legal Agent Benchmark is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does Harvey's Legal Agent Benchmark count toward the BenchLeader Index?
- No. Harvey's Legal Agent Benchmark is shown for reference but left out of the composite index.