SWE Atlas: Test Writing
Writing tests for existing code. Scale AI.
As of 19 Sept 2026, Claude Fable 5.1 leads SWE Atlas: Test Writing on BenchLeader with 67.0%, ahead of Claude Opus 5 at 62.2%, across 22 model configurations with a published result.
- Published by
- Scale AI SEAL
- Category
- Coding
- Index weight
- Reference only
- Models
- 22
- Data as of
- 19 Sept 2026
Scale AI SEAL Leaderboards.
What the test looks like
The model writes tests for real code and is scored on coverage and correctness.
How it is scored
Score, published by Scale AI.
What to keep in mind
A specific engineering skill.
- 1Claude Fable 5.1 (xhigh)67.0%
- 2Claude Opus 5 (xhigh)62.2%
- 3Claude Fable 5 (xhigh)55.6%
- 4Gemini 3.8 Flash53.7%
- 5GPT-6 Astra (xhigh)50.7%
- 6Claude Opus 4.8 (xhigh)49.6%
- 7GPT-5.6 Sol (xhigh)45.9%
- 8GPT-5.4 (xhigh)44.4%
- 9GPT-5.5 (xhigh)42.6%
- 10Muse Spark 1.1 (xhigh)41.5%
- 11GLM-5.241.5%
- 12GPT-5.3 Chat (xhigh)39.0%
- 13Claude Opus 4.738.5%
- 14Claude Opus 4.636.7%
- 15Claude Sonnet 4.631.8%
22 of 22
| # | ||||
|---|---|---|---|---|
| 1 | 67.0% | 71.6 | 2026-09-02 | |
| 2 | 62.2% | 70.2 | 2026-08-04 | |
| 3 | 55.6% | – | 2026-06-11 | |
| 4 | 53.7% | – | 2026-09-09 | |
| 5 | 50.7% | 71.1 | 2026-09-09 | |
| 6 | 49.6% | – | 2026-06-08 | |
| 7 | 45.9% | 68.2 | 2026-07-28 | |
| 8 | 44.4% | 65.4 | 2026-03-26 | |
| 9 | 42.6% | 67.6 | 2026-05-07 | |
| 10 | 41.5% | 61.8 | 2026-07-09 | |
| 11 | 41.5% | 52.1 | 2026-06-23 | |
| 12 | 39.0% | – | 2026-03-26 | |
| 13 | 38.5% | 64.5 | 2026-06-18 | |
| 14 | 36.7% | 63.7 | 2026-03-26 | |
| 15 | 31.8% | 59.0 | 2026-03-26 | |
| 16 | 31.1% | 65.7 | 2026-04-08 | |
| 17 | 30.3% | 57.7 | 2026-03-26 | |
| 18 | 29.8% | 63.9 | 2026-03-26 | |
| 19 | 28.7% | 54.7 | 2026-03-26 | |
| 20 | 27.1% | 52.6 | 2026-06-18 | |
| 21 | 25.8% | 53.9 | 2026-03-26 | |
| 22 | 18.6% | 54.2 | 2026-03-26 |
Cite as: BenchLeader, “SWE Atlas: Test Writing leaderboard”, https://www.benchleader.com/benchmarks/scale_sweatlas_tw, data as of 19 Sept 2026.
SWE Atlas: Test Writing: questions
- What does SWE Atlas: Test Writing measure?
- The model writes tests for real code and is scored on coverage and correctness. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads SWE Atlas: Test Writing?
- Claude Fable 5.1 leads SWE Atlas: Test Writing with 67.0% as of 19 Sept 2026, ahead of Claude Opus 5 at 62.2%.
- How many models have SWE Atlas: Test Writing results?
- 22 model configurations have a SWE Atlas: Test Writing result on BenchLeader, all taken from Scale AI SEAL.
- Who runs SWE Atlas: Test Writing and how often is it updated?
- SWE Atlas: Test Writing is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does SWE Atlas: Test Writing count toward the BenchLeader Index?
- No. SWE Atlas: Test Writing is shown for reference but left out of the composite index.