BenchLeader

CL-bench

Continual-learning style tasks that require using information given in context.

As of 19 Sept 2026, GPT-5.4 leads CL-bench on BenchLeader with 27.9%, ahead of GPT-5.1 at 23.7%, across 21 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Reasoning
Index weight
Reference only
Models
21
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Tasks where the model must apply new rules or facts supplied in the prompt rather than prior knowledge.

How it is scored

Overall score, as published by CL-bench.

What to keep in mind

New; small set.

21 of 21
#
1GPT-5.4xhighOpenAI27.9%65.42026-03-05
2GPT-5.1highOpenAI23.7%58.72025-11-13
3Grok 4.20xAI22.2%2026-02-17
4Claude Opus 4.5Anthropic21.1%58.22025-11-24
5GPT-5.1OpenAI21.1%57.22025-11-13
6Gemini 3.1 ProGoogle20.8%63.92026-02-19
7Claude Opus 4.6Anthropic20.7%63.72026-02-05
8Qwen3.6 PlusAlibabaopen20.3%58.52026-03-31
9Qwen3.5 PlusAlibaba19.8%2026-02-16
10Kimi K2.5Moonshot AIopen19.3%53.92026-01-27
11GLM-5Zhipu AIopen18.7%54.72026-02-11
12GPT-5.2OpenAI18.2%58.42025-12-11
13GPT-5.2highOpenAI18.1%55.72025-12-11
14o3highOpenAI17.8%53.02025-04-16
15Kimi K2 ThinkingthinkingMoonshot AIopen17.6%52.22025-11-06
16GLM-4.7Zhipu AIopen ↗15.9%52.42025-12-22
17Gemini 3 ProGoogle15.8%61.12025-11-18
18Mimo v2 ProXiaomi15.7%59.8
19Qwen3 MaxmaxAlibaba14.5%53.52025-09-24
20DeepSeek V3.2thinkingDeepSeekopen13.2%56.12025-09-29
21MiniMax-M2.5MiniMaxopen ↗11.4%54.22026-02-12

Cite as: BenchLeader, “CL-bench leaderboard”, https://www.benchleader.com/benchmarks/cl_bench, data as of 19 Sept 2026.

CL-bench: questions

What does CL-bench measure?
Tasks where the model must apply new rules or facts supplied in the prompt rather than prior knowledge. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads CL-bench?
GPT-5.4 leads CL-bench with 27.9% as of 19 Sept 2026, ahead of GPT-5.1 at 23.7%.
How many models have CL-bench results?
21 model configurations have a CL-bench result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs CL-bench and how often is it updated?
CL-bench is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does CL-bench count toward the BenchLeader Index?
No. CL-bench is shown for reference but left out of the composite index.