BenchLeader

CaseLaw v2

Questions over Canadian court cases. Run by Vals AI.

As of 19 Sept 2026, Grok 4.3 leads CaseLaw v2 on BenchLeader with 79.3%, ahead of GPT-5.1 at 73.4%, across 52 model configurations with a published result.

Published by
Vals AI
Category
Knowledge
Index weight
Reference only
Models
52
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

Questions about Canadian court decisions that require reading the case and applying it, written by lawyers.

How it is scored

Accuracy, run by Vals AI.

What to keep in mind

Canada-specific; complements the US-centric LegalBench.

52 of 52
#
1Grok 4.3xAI79.3%50.52026-05-04
2GPT-5.1highOpenAI73.4%58.72026-05-04
3GPT-4.1highOpenAI69.9%46.62026-05-04
4GPT-5 minihighOpenAI68.5%55.32026-05-04
5Claude Opus 4.7Anthropic68.4%64.52026-05-04
6GPT-5highOpenAI66.5%58.62026-05-04
7GPT-5.5xhighOpenAI66.2%67.62026-05-04
8GPT-5.2xhighOpenAI66.0%62.02026-05-04
9Grok 4xAI65.8%58.52026-05-04
10Grok 4 FastthinkingxAI65.7%52.62026-05-04
11Kimi K2 ThinkingthinkingMoonshot AIopen65.7%52.22026-05-04
12Gemini 3.1 ProhighGoogle64.8%60.72026-05-04
13Command ACohereopen64.5%42.32026-05-04
14Claude Sonnet 4.6Anthropic64.0%59.02026-05-04
15Gemini 2.5 ProGoogle63.9%54.52026-05-04
16GPT-5.4xhighOpenAI63.8%65.42026-05-04
17Muse SparkMeta63.1%65.72026-05-04
18Claude Opus 4.5thinkingAnthropic62.6%60.82026-05-04
19Claude Sonnet 4.5thinkingAnthropic62.2%53.42026-05-04
20Claude Opus 4.6thinkingAnthropic62.1%62.52026-05-04
21Mistral Large 3Mistral AIopen61.4%46.12026-05-04
22Kimi K2.6Moonshot AIopen ↗61.2%60.52026-05-04
23MiniMax-M2.7MiniMaxopen ↗60.9%56.12026-05-04
24Grok 4.1thinkingxAI60.5%56.02026-05-04
25Qwen3.5 PlusthinkingAlibaba59.7%57.32026-05-04
26GPT-4ohighOpenAI59.7%41.62026-05-04
27DeepSeek V4 PromaxDeepSeekopen ↗59.4%64.22026-05-04
28Kimi K2.5thinkingMoonshot AIopen58.7%58.52026-05-04
29Trinity LargethinkingArcee AIopen57.9%47.62026-05-04
30Claude Haiku 4.5thinkingAnthropic56.5%47.52026-05-04
31Qwen3.5 FlashAlibaba56.0%52.52026-05-04
32Gemini 3 FlashhighGoogle55.8%58.72026-05-04
33MiniMax-M2MiniMaxopen55.8%55.12026-05-04
34DeepSeek V3.2highDeepSeekopen55.4%2026-05-04
35Qwen3 MaxmaxAlibaba55.0%53.52026-05-04
36Gemini 3.1 Flash LitehighGoogle55.0%49.12026-05-04
37GLM-4.7Zhipu AIopen ↗54.9%52.42026-05-04
38Grok 4.20thinkingxAI54.5%59.92026-05-04
39MiniMax-M2.5MiniMaxopen ↗53.5%54.22026-05-04
40Qwen3.6 27BAlibaba53.2%44.72026-05-04
41Gemini 3 ProhighGoogle53.1%61.12026-05-04
42GPT-5 nanohighOpenAI52.6%48.22026-05-04
43Gemma 4 31BhighGoogleopen ↗52.6%2026-05-04
44GLM-5thinkingZhipu AIopen52.5%59.62026-05-04
45GPT-5.4 nanohighOpenAI51.9%49.12026-05-04
46GPT-5.4 minixhighOpenAI51.7%55.72026-05-04
47GLM-5.1Zhipu AIopen ↗51.5%56.92026-05-04
48Qwen3.6 PlusAlibabaopen51.5%58.52026-05-04
49gpt-oss-120bOpenAIopen48.8%46.92026-05-04
50Qwen3 6maxAlibaba47.9%63.32026-05-04
51Mistral Medium 3.5highMistral AIopen44.2%2026-05-04
52gpt-oss-20bOpenAIopen43.8%44.32026-05-04

Cite as: BenchLeader, “CaseLaw v2 leaderboard”, https://www.benchleader.com/benchmarks/vals_caselaw, data as of 19 Sept 2026.

CaseLaw v2: questions

What does CaseLaw v2 measure?
Questions about Canadian court decisions that require reading the case and applying it, written by lawyers. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads CaseLaw v2?
Grok 4.3 leads CaseLaw v2 with 79.3% as of 19 Sept 2026, ahead of GPT-5.1 at 73.4%.
How many models have CaseLaw v2 results?
52 model configurations have a CaseLaw v2 result on BenchLeader, all taken from Vals AI.
Who runs CaseLaw v2 and how often is it updated?
CaseLaw v2 is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does CaseLaw v2 count toward the BenchLeader Index?
No. CaseLaw v2 is shown for reference but left out of the composite index.