PostTrainBench
Post-training a small model to a target, end to end.
As of 19 Sept 2026, Claude Fable 5 leads PostTrainBench on BenchLeader with 41.8%, ahead of GPT-5.6 Sol at 36.2%, across 12 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 12
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
The agent must run a post-training pipeline on a small model and hit target metrics.
How it is scored
Average score, as published.
What to keep in mind
Very few models measured; a research-engineering task.
- 1Claude Fable 5 (max)41.8%
- 2GPT-5.6 Sol (max)36.2%
- 3Claude Opus 535.0%
- 4Claude Opus 4.8 (high)33.8%
- 5Claude Opus 4.8 (max)32.9%
- 6Kimi K332.0%
- 7GLM-5.2 (max)31.7%
- 8Claude Opus 4.7 (xhigh)28.6%
- 9GPT-5.5 (xhigh)27.2%
- 10Grok 4.5 (high)23.4%
- 11Gemini 3.1 Pro22.0%
- 12GPT-5.4 (high)19.0%
12 of 12
| # | ||||
|---|---|---|---|---|
| 1 | 41.8% | 65.8 | 2026-06-09 | |
| 2 | 36.2% | 68.8 | 2026-07-09 | |
| 3 | 35.0% | 66.7 | 2026-07-24 | |
| 4 | 33.8% | 62.2 | 2026-05-28 | |
| 5 | 32.9% | 64.2 | 2026-05-28 | |
| 6 | 32.0% | 64.2 | 2026-07-16 | |
| 7 | 31.7% | 63.9 | 2026-06-16 | |
| 8 | 28.6% | 57.3 | 2026-04-16 | |
| 9 | 27.2% | 67.6 | 2026-04-23 | |
| 10 | 23.4% | 61.0 | 2026-07-08 | |
| 11 | 22.0% | 63.9 | 2026-02-19 | |
| 12 | 19.0% | 59.0 | 2026-03-05 |
Cite as: BenchLeader, “PostTrainBench leaderboard”, https://www.benchleader.com/benchmarks/posttrainbench, data as of 19 Sept 2026.
PostTrainBench: questions
- What does PostTrainBench measure?
- The agent must run a post-training pipeline on a small model and hit target metrics. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads PostTrainBench?
- Claude Fable 5 leads PostTrainBench with 41.8% as of 19 Sept 2026, ahead of GPT-5.6 Sol at 36.2%.
- How many models have PostTrainBench results?
- 12 model configurations have a PostTrainBench result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs PostTrainBench and how often is it updated?
- PostTrainBench is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does PostTrainBench count toward the BenchLeader Index?
- No. PostTrainBench is shown for reference but left out of the composite index.