BenchLeader

SWE Atlas: Test Writing

Writing tests for existing code. Scale AI.

As of 19 Sept 2026, Claude Fable 5.1 leads SWE Atlas: Test Writing on BenchLeader with 67.0%, ahead of Claude Opus 5 at 62.2%, across 22 model configurations with a published result.

Published by
Scale AI SEAL
Category
Coding
Index weight
Reference only
Models
22
Data as of
19 Sept 2026

Scale AI SEAL Leaderboards.

What the test looks like

The model writes tests for real code and is scored on coverage and correctness.

How it is scored

Score, published by Scale AI.

What to keep in mind

A specific engineering skill.

22 of 22
#
1Claude Fable 5.1xhighAnthropic67.0%71.62026-09-02
2Claude Opus 5xhighAnthropic62.2%70.22026-08-04
3Claude Fable 5xhighAnthropic55.6%2026-06-11
4Gemini 3.8 FlashGoogle53.7%2026-09-09
5GPT-6 AstraxhighOpenAI50.7%71.12026-09-09
6Claude Opus 4.8xhighAnthropic49.6%2026-06-08
7GPT-5.6 SolxhighOpenAI45.9%68.22026-07-28
8GPT-5.4xhighOpenAI44.4%65.42026-03-26
9GPT-5.5xhighOpenAI42.6%67.62026-05-07
10Muse Spark 1.1xhighMeta41.5%61.82026-07-09
11GLM-5.2Zhipu AIopen ↗41.5%52.12026-06-23
12GPT-5.3 ChatxhighOpenAI39.0%2026-03-26
13Claude Opus 4.7Anthropic38.5%64.52026-06-18
14Claude Opus 4.6Anthropic36.7%63.72026-03-26
15Claude Sonnet 4.6Anthropic31.8%59.02026-03-26
16Muse SparkMeta31.1%65.72026-04-08
17Gemini 3 FlashGoogle30.3%57.72026-03-26
18Gemini 3.1 ProGoogle29.8%63.92026-03-26
19GLM-5Zhipu AIopen28.7%54.72026-03-26
20DeepSeek V4 ProDeepSeekopen ↗27.1%52.62026-06-18
21Kimi K2.5Moonshot AIopen25.8%53.92026-03-26
22MiniMax-M2.5MiniMaxopen ↗18.6%54.22026-03-26

Cite as: BenchLeader, “SWE Atlas: Test Writing leaderboard”, https://www.benchleader.com/benchmarks/scale_sweatlas_tw, data as of 19 Sept 2026.

SWE Atlas: Test Writing: questions

What does SWE Atlas: Test Writing measure?
The model writes tests for real code and is scored on coverage and correctness. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads SWE Atlas: Test Writing?
Claude Fable 5.1 leads SWE Atlas: Test Writing with 67.0% as of 19 Sept 2026, ahead of Claude Opus 5 at 62.2%.
How many models have SWE Atlas: Test Writing results?
22 model configurations have a SWE Atlas: Test Writing result on BenchLeader, all taken from Scale AI SEAL.
Who runs SWE Atlas: Test Writing and how often is it updated?
SWE Atlas: Test Writing is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
Does SWE Atlas: Test Writing count toward the BenchLeader Index?
No. SWE Atlas: Test Writing is shown for reference but left out of the composite index.