BenchLeader

SkillsBench

How much reusable skills help an agent complete tasks. Run by Vals AI.

As of 19 Sept 2026, DeepSeek V4.1 Flash leads SkillsBench on BenchLeader with 69.8%, ahead of Grok 4.5 at 66.0%, across 34 model configurations with a published result.

Published by
Vals AI
Category
Agents & tools
Index weight
Reference only
Models
34
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

Agents complete tasks with and without a library of reusable skills; the score reflects task success with skills available.

How it is scored

Percent of tasks completed, run by Vals AI.

What to keep in mind

Measures a specific agent design choice; not a general capability test.

34 of 34
#
1DeepSeek V4.1 FlashhighDeepSeekopen ↗69.8%2026-09-11
2Grok 4.5highxAI66.0%61.02026-09-11
3Gemini 3.7 FlashhighGoogle65.9%64.62026-09-11
4GPT 5.5 CodexOpenAI62.6%2026-09-11
5GPT-5.5xhighOpenAI62.2%67.62026-09-11
6Claude Fable 5.1Anthropic61.5%64.52026-09-11
7GPT-5.6 LunamaxOpenAI60.5%60.22026-09-11
8Claude Opus 5Anthropic60.4%66.72026-09-11
9Claude Opus 4.8Anthropic59.2%61.92026-09-11
10Muse Spark 1.1xhighMeta59.2%61.82026-09-11
11GPT-5.6 TerramaxOpenAI58.9%65.02026-09-11
12Gemini 3.8 FlashhighGoogle58.0%64.52026-09-11
13Grok 4.6highxAI55.8%64.52026-09-11
14Qwen3.7 PlusAlibaba54.3%59.92026-09-11
15GPT-5.6 SolmaxOpenAI54.1%68.82026-09-11
16DeepSeek V4 Pro 0813maxDeepSeekopen ↗53.8%56.02026-09-11
17Muse Spark 1.2xhighMeta53.0%64.12026-09-11
18Gemini 3.5 FlashhighGoogle52.7%63.62026-09-11
19GPT-5.4xhighOpenAI51.7%65.42026-09-11
20MiniMax-M3MiniMaxopen ↗51.5%57.52026-09-11
21DeepSeek V4 PromaxDeepSeekopen ↗51.3%64.22026-09-11
22DeepSeek V4 Flash 0731highDeepSeekopen ↗50.7%55.02026-09-11
23Kimi K2.7 CodeMoonshot AIopen ↗50.0%55.92026-09-11
24Claude Sonnet 4.6Anthropic49.0%59.02026-09-11
25GLM-5.3maxZhipu AIopen ↗47.5%65.82026-09-11
26Claude Sonnet 5Anthropic46.5%57.22026-09-11
27GLM-5.2maxZhipu AIopen ↗45.1%63.92026-09-11
28Qwen3 8maxAlibaba42.0%66.52026-09-11
29Grok 4.3highxAI40.6%58.22026-09-11
30GLM-5.3-FlashmaxZhipu AIopen ↗40.2%55.32026-09-11
31Qwen3.8 27BxhighAlibabaopen ↗38.1%59.22026-09-11
32Inkling SmallThinking Machinesopen ↗33.6%56.32026-09-11
33Ling 3.0 FlashAnt Groupopen ↗27.9%53.62026-09-11
34Mercury 2.5highInception18.1%2026-09-11

Cite as: BenchLeader, “SkillsBench leaderboard”, https://www.benchleader.com/benchmarks/vals_skillsbench, data as of 19 Sept 2026.

SkillsBench: questions

What does SkillsBench measure?
Agents complete tasks with and without a library of reusable skills; the score reflects task success with skills available. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads SkillsBench?
DeepSeek V4.1 Flash leads SkillsBench with 69.8% as of 19 Sept 2026, ahead of Grok 4.5 at 66.0%.
How many models have SkillsBench results?
34 model configurations have a SkillsBench result on BenchLeader, all taken from Vals AI.
Who runs SkillsBench and how often is it updated?
SkillsBench is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does SkillsBench count toward the BenchLeader Index?
No. SkillsBench is shown for reference but left out of the composite index.