CL-bench
Continual-learning style tasks that require using information given in context.
As of 19 Sept 2026, GPT-5.4 leads CL-bench on BenchLeader with 27.9%, ahead of GPT-5.1 at 23.7%, across 21 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Reasoning
- Index weight
- Reference only
- Models
- 21
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Tasks where the model must apply new rules or facts supplied in the prompt rather than prior knowledge.
How it is scored
Overall score, as published by CL-bench.
What to keep in mind
New; small set.
- 1GPT-5.4 (xhigh)27.9%
- 2GPT-5.1 (high)23.7%
- 3Grok 4.2022.2%
- 4Claude Opus 4.521.1%
- 5GPT-5.121.1%
- 6Gemini 3.1 Pro20.8%
- 7Claude Opus 4.620.7%
- 8Qwen3.6 Plus20.3%
- 9Qwen3.5 Plus19.8%
- 10Kimi K2.519.3%
- 11GLM-518.7%
- 12GPT-5.218.2%
- 13GPT-5.2 (high)18.1%
- 14o3 (high)17.8%
- 15Kimi K2 Thinking (thinking)17.6%
21 of 21
| # | ||||
|---|---|---|---|---|
| 1 | 27.9% | 65.4 | 2026-03-05 | |
| 2 | 23.7% | 58.7 | 2025-11-13 | |
| 3 | 22.2% | – | 2026-02-17 | |
| 4 | 21.1% | 58.2 | 2025-11-24 | |
| 5 | 21.1% | 57.2 | 2025-11-13 | |
| 6 | 20.8% | 63.9 | 2026-02-19 | |
| 7 | 20.7% | 63.7 | 2026-02-05 | |
| 8 | 20.3% | 58.5 | 2026-03-31 | |
| 9 | 19.8% | – | 2026-02-16 | |
| 10 | 19.3% | 53.9 | 2026-01-27 | |
| 11 | 18.7% | 54.7 | 2026-02-11 | |
| 12 | 18.2% | 58.4 | 2025-12-11 | |
| 13 | 18.1% | 55.7 | 2025-12-11 | |
| 14 | 17.8% | 53.0 | 2025-04-16 | |
| 15 | 17.6% | 52.2 | 2025-11-06 | |
| 16 | 15.9% | 52.4 | 2025-12-22 | |
| 17 | 15.8% | 61.1 | 2025-11-18 | |
| 18 | 15.7% | 59.8 | – | |
| 19 | 14.5% | 53.5 | 2025-09-24 | |
| 20 | 13.2% | 56.1 | 2025-09-29 | |
| 21 | 11.4% | 54.2 | 2026-02-12 |
Cite as: BenchLeader, “CL-bench leaderboard”, https://www.benchleader.com/benchmarks/cl_bench, data as of 19 Sept 2026.
CL-bench: questions
- What does CL-bench measure?
- Tasks where the model must apply new rules or facts supplied in the prompt rather than prior knowledge. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads CL-bench?
- GPT-5.4 leads CL-bench with 27.9% as of 19 Sept 2026, ahead of GPT-5.1 at 23.7%.
- How many models have CL-bench results?
- 21 model configurations have a CL-bench result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs CL-bench and how often is it updated?
- CL-bench is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does CL-bench count toward the BenchLeader Index?
- No. CL-bench is shown for reference but left out of the composite index.