BenchLeader

SWE Atlas: Refactoring

Refactoring tasks in real repositories. Scale AI.

As of 19 Sept 2026, GPT-6 Astra leads SWE Atlas: Refactoring on BenchLeader with 59.0%, ahead of Claude Fable 5.1 at 56.7%, across 17 model configurations with a published result.

Published by
Scale AI SEAL
Category
Coding
Index weight
Reference only
Models
17
Data as of
19 Sept 2026

Scale AI SEAL Leaderboards.

What the test looks like

Multi-file refactors in real repositories, checked by tests and review criteria.

How it is scored

Percent of refactors accepted, published by Scale AI.

What to keep in mind

Harness-dependent, like all agentic coding boards.

17 of 17
#
1GPT-6 AstraxhighOpenAI59.0%71.12026-09-09
2Claude Fable 5.1xhighAnthropic56.7%71.62026-09-02
3Claude Fable 5xhighAnthropic54.8%2026-06-11
4Claude Opus 4.7Anthropic48.6%64.52026-05-06
5Claude Opus 4.8Anthropic46.7%61.92026-06-08
6GPT-5.5xhighOpenAI44.8%67.62026-05-06
7Gemini 3.8 FlashGoogle44.8%2026-09-09
8GPT-5.4xhighOpenAI44.3%65.42026-05-06
9GLM-5.2Zhipu AIopen ↗42.4%52.12026-06-23
10GPT-5.3 ChatxhighOpenAI42.4%2026-05-06
11Claude Opus 4.6Anthropic35.6%63.72026-05-06
12Gemini 3.1 ProGoogle33.8%63.92026-05-06
13Claude Sonnet 4.6Anthropic32.2%59.02026-05-06
14GLM-5Zhipu AIopen24.2%54.72026-05-06
15Kimi K2.5Moonshot AIopen20.9%53.92026-05-06
16MiniMax-M2.5MiniMaxopen ↗19.5%54.22026-05-06
17Gemini 3 FlashGoogle10.0%57.72026-05-06

Cite as: BenchLeader, “SWE Atlas: Refactoring leaderboard”, https://www.benchleader.com/benchmarks/scale_sweatlas_refactoring, data as of 19 Sept 2026.

SWE Atlas: Refactoring: questions

What does SWE Atlas: Refactoring measure?
Multi-file refactors in real repositories, checked by tests and review criteria. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads SWE Atlas: Refactoring?
GPT-6 Astra leads SWE Atlas: Refactoring with 59.0% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 56.7%.
How many models have SWE Atlas: Refactoring results?
17 model configurations have a SWE Atlas: Refactoring result on BenchLeader, all taken from Scale AI SEAL.
Who runs SWE Atlas: Refactoring and how often is it updated?
SWE Atlas: Refactoring is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
Does SWE Atlas: Refactoring count toward the BenchLeader Index?
No. SWE Atlas: Refactoring is shown for reference but left out of the composite index.