BenchLeader

Public Benefits Bench

Navigating SNAP benefits questions for applicants. Run by Vals AI.

As of 19 Sept 2026, Claude Opus 5 leads Public Benefits Bench on BenchLeader with 76.9%, ahead of Claude Fable 5.1 at 74.9%, across 34 model configurations with a published result.

Published by
Vals AI
Category
Knowledge
Index weight
Reference only
Models
34
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

Questions a benefits applicant might ask about SNAP eligibility and process, graded against policy.

How it is scored

Accuracy, run by Vals AI.

What to keep in mind

US-specific and policy changes over time.

34 of 34
#
1Claude Opus 5Anthropic76.9%66.72026-09-16
2Claude Fable 5.1Anthropic74.9%64.52026-09-16
3Claude Fable 5Anthropic70.4%68.32026-09-16
4GLM-5.3maxZhipu AIopen ↗68.5%65.82026-09-16
5Muse Spark 1.2xhighMeta68.5%64.12026-09-16
6Kimi K3Moonshot AIopen ↗68.3%64.22026-09-16
7Claude Opus 4.8Anthropic68.1%61.92026-09-16
8Qwen3 8maxAlibaba67.1%66.52026-09-16
9Grok 4.6highxAI66.8%64.52026-09-16
10GPT-5.6 SolmaxOpenAI66.5%68.82026-09-16
11Claude Sonnet 5Anthropic66.0%57.22026-09-16
12Gemini 3.8 FlashhighGoogle65.3%64.52026-09-16
13DeepSeek V4.1 FlashhighDeepSeekopen ↗64.3%2026-09-16
14MiniMax-M3MiniMaxopen ↗64.1%57.52026-09-16
15DeepSeek V4 PromaxDeepSeekopen ↗62.9%64.22026-09-16
16Claude Sonnet 4.6Anthropic62.5%59.02026-09-16
17GPT-5.6 TerraxhighOpenAI62.4%64.32026-09-16
18GLM-5.1Zhipu AIopen ↗61.8%56.92026-09-16
19Muse Spark 1.1xhighMeta61.3%61.82026-09-16
20GPT-5.6 LunamaxOpenAI61.2%60.22026-09-16
21GPT-5.5xhighOpenAI60.9%67.62026-09-16
22Gemini 3.5 FlashhighGoogle59.5%63.62026-09-16
23Inkling SmallThinking Machinesopen ↗59.4%56.32026-09-16
24Kimi K2.6Moonshot AIopen ↗56.6%60.52026-09-16
25Gemini 3.6 FlashhighGoogle56.6%61.72026-09-16
26Claude Haiku 4.5Anthropic54.3%48.12026-09-16
27Gemini 3.1 ProhighGoogle53.8%60.72026-09-16
28Gemini 3.5 Flash LitehighGoogle51.8%44.32026-09-16
29Grok 4.3highxAI51.7%58.22026-09-16
30Ling 3.0 FlashAnt Groupopen ↗50.7%53.62026-09-16
31Mercury 2.5highInception45.3%2026-09-16
32Grok 4.1thinkingxAI44.8%56.02026-09-16
33Laguna M.1Poolsideopen43.8%37.72026-09-16
34Laguna XS.2Poolside41.4%36.52026-09-16

Cite as: BenchLeader, “Public Benefits Bench leaderboard”, https://www.benchleader.com/benchmarks/vals_public_benefits, data as of 19 Sept 2026.

Public Benefits Bench: questions

What does Public Benefits Bench measure?
Questions a benefits applicant might ask about SNAP eligibility and process, graded against policy. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads Public Benefits Bench?
Claude Opus 5 leads Public Benefits Bench with 76.9% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 74.9%.
How many models have Public Benefits Bench results?
34 model configurations have a Public Benefits Bench result on BenchLeader, all taken from Vals AI.
Who runs Public Benefits Bench and how often is it updated?
Public Benefits Bench is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does Public Benefits Bench count toward the BenchLeader Index?
No. Public Benefits Bench is shown for reference but left out of the composite index.