SWE Atlas: Refactoring
Refactoring tasks in real repositories. Scale AI.
As of 19 Sept 2026, GPT-6 Astra leads SWE Atlas: Refactoring on BenchLeader with 59.0%, ahead of Claude Fable 5.1 at 56.7%, across 17 model configurations with a published result.
- Published by
- Scale AI SEAL
- Category
- Coding
- Index weight
- Reference only
- Models
- 17
- Data as of
- 19 Sept 2026
Scale AI SEAL Leaderboards.
What the test looks like
Multi-file refactors in real repositories, checked by tests and review criteria.
How it is scored
Percent of refactors accepted, published by Scale AI.
What to keep in mind
Harness-dependent, like all agentic coding boards.
- 1GPT-6 Astra (xhigh)59.0%
- 2Claude Fable 5.1 (xhigh)56.7%
- 3Claude Fable 5 (xhigh)54.8%
- 4Claude Opus 4.748.6%
- 5Claude Opus 4.846.7%
- 6GPT-5.5 (xhigh)44.8%
- 7Gemini 3.8 Flash44.8%
- 8GPT-5.4 (xhigh)44.3%
- 9GLM-5.242.4%
- 10GPT-5.3 Chat (xhigh)42.4%
- 11Claude Opus 4.635.6%
- 12Gemini 3.1 Pro33.8%
- 13Claude Sonnet 4.632.2%
- 14GLM-524.2%
- 15Kimi K2.520.9%
17 of 17
| # | ||||
|---|---|---|---|---|
| 1 | 59.0% | 71.1 | 2026-09-09 | |
| 2 | 56.7% | 71.6 | 2026-09-02 | |
| 3 | 54.8% | – | 2026-06-11 | |
| 4 | 48.6% | 64.5 | 2026-05-06 | |
| 5 | 46.7% | 61.9 | 2026-06-08 | |
| 6 | 44.8% | 67.6 | 2026-05-06 | |
| 7 | 44.8% | – | 2026-09-09 | |
| 8 | 44.3% | 65.4 | 2026-05-06 | |
| 9 | 42.4% | 52.1 | 2026-06-23 | |
| 10 | 42.4% | – | 2026-05-06 | |
| 11 | 35.6% | 63.7 | 2026-05-06 | |
| 12 | 33.8% | 63.9 | 2026-05-06 | |
| 13 | 32.2% | 59.0 | 2026-05-06 | |
| 14 | 24.2% | 54.7 | 2026-05-06 | |
| 15 | 20.9% | 53.9 | 2026-05-06 | |
| 16 | 19.5% | 54.2 | 2026-05-06 | |
| 17 | 10.0% | 57.7 | 2026-05-06 |
Cite as: BenchLeader, “SWE Atlas: Refactoring leaderboard”, https://www.benchleader.com/benchmarks/scale_sweatlas_refactoring, data as of 19 Sept 2026.
SWE Atlas: Refactoring: questions
- What does SWE Atlas: Refactoring measure?
- Multi-file refactors in real repositories, checked by tests and review criteria. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads SWE Atlas: Refactoring?
- GPT-6 Astra leads SWE Atlas: Refactoring with 59.0% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 56.7%.
- How many models have SWE Atlas: Refactoring results?
- 17 model configurations have a SWE Atlas: Refactoring result on BenchLeader, all taken from Scale AI SEAL.
- Who runs SWE Atlas: Refactoring and how often is it updated?
- SWE Atlas: Refactoring is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does SWE Atlas: Refactoring count toward the BenchLeader Index?
- No. SWE Atlas: Refactoring is shown for reference but left out of the composite index.