CaseLaw v2
Questions over Canadian court cases. Run by Vals AI.
As of 19 Sept 2026, Grok 4.3 leads CaseLaw v2 on BenchLeader with 79.3%, ahead of GPT-5.1 at 73.4%, across 52 model configurations with a published result.
- Published by
- Vals AI
- Category
- Knowledge
- Index weight
- Reference only
- Models
- 52
- Data as of
- 19 Sept 2026
Vals AI (vals.ai).
What the test looks like
Questions about Canadian court decisions that require reading the case and applying it, written by lawyers.
How it is scored
Accuracy, run by Vals AI.
What to keep in mind
Canada-specific; complements the US-centric LegalBench.
- 1Grok 4.379.3%
- 2GPT-5.1 (high)73.4%
- 3GPT-4.1 (high)69.9%
- 4GPT-5 mini (high)68.5%
- 5Claude Opus 4.768.4%
- 6GPT-5 (high)66.5%
- 7GPT-5.5 (xhigh)66.2%
- 8GPT-5.2 (xhigh)66.0%
- 9Grok 465.8%
- 10Grok 4 Fast (thinking)65.7%
- 11Kimi K2 Thinking (thinking)65.7%
- 12Gemini 3.1 Pro (high)64.8%
- 13Command A64.5%
- 14Claude Sonnet 4.664.0%
- 15Gemini 2.5 Pro63.9%
52 of 52
| # | ||||
|---|---|---|---|---|
| 1 | 79.3% | 50.5 | 2026-05-04 | |
| 2 | 73.4% | 58.7 | 2026-05-04 | |
| 3 | 69.9% | 46.6 | 2026-05-04 | |
| 4 | 68.5% | 55.3 | 2026-05-04 | |
| 5 | 68.4% | 64.5 | 2026-05-04 | |
| 6 | 66.5% | 58.6 | 2026-05-04 | |
| 7 | 66.2% | 67.6 | 2026-05-04 | |
| 8 | 66.0% | 62.0 | 2026-05-04 | |
| 9 | 65.8% | 58.5 | 2026-05-04 | |
| 10 | 65.7% | 52.6 | 2026-05-04 | |
| 11 | 65.7% | 52.2 | 2026-05-04 | |
| 12 | 64.8% | 60.7 | 2026-05-04 | |
| 13 | 64.5% | 42.3 | 2026-05-04 | |
| 14 | 64.0% | 59.0 | 2026-05-04 | |
| 15 | 63.9% | 54.5 | 2026-05-04 | |
| 16 | 63.8% | 65.4 | 2026-05-04 | |
| 17 | 63.1% | 65.7 | 2026-05-04 | |
| 18 | 62.6% | 60.8 | 2026-05-04 | |
| 19 | 62.2% | 53.4 | 2026-05-04 | |
| 20 | 62.1% | 62.5 | 2026-05-04 | |
| 21 | 61.4% | 46.1 | 2026-05-04 | |
| 22 | 61.2% | 60.5 | 2026-05-04 | |
| 23 | 60.9% | 56.1 | 2026-05-04 | |
| 24 | 60.5% | 56.0 | 2026-05-04 | |
| 25 | 59.7% | 57.3 | 2026-05-04 | |
| 26 | 59.7% | 41.6 | 2026-05-04 | |
| 27 | 59.4% | 64.2 | 2026-05-04 | |
| 28 | 58.7% | 58.5 | 2026-05-04 | |
| 29 | 57.9% | 47.6 | 2026-05-04 | |
| 30 | 56.5% | 47.5 | 2026-05-04 | |
| 31 | 56.0% | 52.5 | 2026-05-04 | |
| 32 | 55.8% | 58.7 | 2026-05-04 | |
| 33 | 55.8% | 55.1 | 2026-05-04 | |
| 34 | 55.4% | – | 2026-05-04 | |
| 35 | 55.0% | 53.5 | 2026-05-04 | |
| 36 | 55.0% | 49.1 | 2026-05-04 | |
| 37 | 54.9% | 52.4 | 2026-05-04 | |
| 38 | 54.5% | 59.9 | 2026-05-04 | |
| 39 | 53.5% | 54.2 | 2026-05-04 | |
| 40 | 53.2% | 44.7 | 2026-05-04 | |
| 41 | 53.1% | 61.1 | 2026-05-04 | |
| 42 | 52.6% | 48.2 | 2026-05-04 | |
| 43 | 52.6% | – | 2026-05-04 | |
| 44 | 52.5% | 59.6 | 2026-05-04 | |
| 45 | 51.9% | 49.1 | 2026-05-04 | |
| 46 | 51.7% | 55.7 | 2026-05-04 | |
| 47 | 51.5% | 56.9 | 2026-05-04 | |
| 48 | 51.5% | 58.5 | 2026-05-04 | |
| 49 | 48.8% | 46.9 | 2026-05-04 | |
| 50 | 47.9% | 63.3 | 2026-05-04 | |
| 51 | 44.2% | – | 2026-05-04 | |
| 52 | 43.8% | 44.3 | 2026-05-04 |
Cite as: BenchLeader, “CaseLaw v2 leaderboard”, https://www.benchleader.com/benchmarks/vals_caselaw, data as of 19 Sept 2026.
CaseLaw v2: questions
- What does CaseLaw v2 measure?
- Questions about Canadian court decisions that require reading the case and applying it, written by lawyers. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads CaseLaw v2?
- Grok 4.3 leads CaseLaw v2 with 79.3% as of 19 Sept 2026, ahead of GPT-5.1 at 73.4%.
- How many models have CaseLaw v2 results?
- 52 model configurations have a CaseLaw v2 result on BenchLeader, all taken from Vals AI.
- Who runs CaseLaw v2 and how often is it updated?
- CaseLaw v2 is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does CaseLaw v2 count toward the BenchLeader Index?
- No. CaseLaw v2 is shown for reference but left out of the composite index.