SkillsBench
How much reusable skills help an agent complete tasks. Run by Vals AI.
As of 19 Sept 2026, DeepSeek V4.1 Flash leads SkillsBench on BenchLeader with 69.8%, ahead of Grok 4.5 at 66.0%, across 34 model configurations with a published result.
- Published by
- Vals AI
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 34
- Data as of
- 19 Sept 2026
Vals AI (vals.ai).
What the test looks like
Agents complete tasks with and without a library of reusable skills; the score reflects task success with skills available.
How it is scored
Percent of tasks completed, run by Vals AI.
What to keep in mind
Measures a specific agent design choice; not a general capability test.
- 1DeepSeek V4.1 Flash (high)69.8%
- 2Grok 4.5 (high)66.0%
- 3Gemini 3.7 Flash (high)65.9%
- 4GPT 5.5 Codex62.6%
- 5GPT-5.5 (xhigh)62.2%
- 6Claude Fable 5.161.5%
- 7GPT-5.6 Luna (max)60.5%
- 8Claude Opus 560.4%
- 9Claude Opus 4.859.2%
- 10Muse Spark 1.1 (xhigh)59.2%
- 11GPT-5.6 Terra (max)58.9%
- 12Gemini 3.8 Flash (high)58.0%
- 13Grok 4.6 (high)55.8%
- 14Qwen3.7 Plus54.3%
- 15GPT-5.6 Sol (max)54.1%
34 of 34
| # | ||||
|---|---|---|---|---|
| 1 | 69.8% | – | 2026-09-11 | |
| 2 | 66.0% | 61.0 | 2026-09-11 | |
| 3 | 65.9% | 64.6 | 2026-09-11 | |
| 4 | 62.6% | – | 2026-09-11 | |
| 5 | 62.2% | 67.6 | 2026-09-11 | |
| 6 | 61.5% | 64.5 | 2026-09-11 | |
| 7 | 60.5% | 60.2 | 2026-09-11 | |
| 8 | 60.4% | 66.7 | 2026-09-11 | |
| 9 | 59.2% | 61.9 | 2026-09-11 | |
| 10 | 59.2% | 61.8 | 2026-09-11 | |
| 11 | 58.9% | 65.0 | 2026-09-11 | |
| 12 | 58.0% | 64.5 | 2026-09-11 | |
| 13 | 55.8% | 64.5 | 2026-09-11 | |
| 14 | 54.3% | 59.9 | 2026-09-11 | |
| 15 | 54.1% | 68.8 | 2026-09-11 | |
| 16 | 53.8% | 56.0 | 2026-09-11 | |
| 17 | 53.0% | 64.1 | 2026-09-11 | |
| 18 | 52.7% | 63.6 | 2026-09-11 | |
| 19 | 51.7% | 65.4 | 2026-09-11 | |
| 20 | 51.5% | 57.5 | 2026-09-11 | |
| 21 | 51.3% | 64.2 | 2026-09-11 | |
| 22 | 50.7% | 55.0 | 2026-09-11 | |
| 23 | 50.0% | 55.9 | 2026-09-11 | |
| 24 | 49.0% | 59.0 | 2026-09-11 | |
| 25 | 47.5% | 65.8 | 2026-09-11 | |
| 26 | 46.5% | 57.2 | 2026-09-11 | |
| 27 | 45.1% | 63.9 | 2026-09-11 | |
| 28 | 42.0% | 66.5 | 2026-09-11 | |
| 29 | 40.6% | 58.2 | 2026-09-11 | |
| 30 | 40.2% | 55.3 | 2026-09-11 | |
| 31 | 38.1% | 59.2 | 2026-09-11 | |
| 32 | Inkling SmallThinking Machinesopen ↗ | 33.6% | 56.3 | 2026-09-11 |
| 33 | 27.9% | 53.6 | 2026-09-11 | |
| 34 | 18.1% | – | 2026-09-11 |
Cite as: BenchLeader, “SkillsBench leaderboard”, https://www.benchleader.com/benchmarks/vals_skillsbench, data as of 19 Sept 2026.
SkillsBench: questions
- What does SkillsBench measure?
- Agents complete tasks with and without a library of reusable skills; the score reflects task success with skills available. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads SkillsBench?
- DeepSeek V4.1 Flash leads SkillsBench with 69.8% as of 19 Sept 2026, ahead of Grok 4.5 at 66.0%.
- How many models have SkillsBench results?
- 34 model configurations have a SkillsBench result on BenchLeader, all taken from Vals AI.
- Who runs SkillsBench and how often is it updated?
- SkillsBench is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does SkillsBench count toward the BenchLeader Index?
- No. SkillsBench is shown for reference but left out of the composite index.