SWE Atlas: Codebase QnA
Answering questions about a large codebase. Scale AI.
As of 19 Sept 2026, Claude Opus 5 leads SWE Atlas: Codebase QnA on BenchLeader with 63.2%, ahead of Claude Fable 5.1 at 60.0%, across 22 model configurations with a published result.
- Published by
- Scale AI SEAL
- Category
- Coding
- Index weight
- Reference only
- Models
- 22
- Data as of
- 19 Sept 2026
Scale AI SEAL Leaderboards.
What the test looks like
The model must answer questions about how a real codebase works, requiring navigation and reading across files.
How it is scored
Accuracy, published by Scale AI.
What to keep in mind
Measures code comprehension rather than editing.
- 1Claude Opus 5 (xhigh)63.2%
- 2Claude Fable 5.1 (xhigh)60.0%
- 3GPT-6 Astra (xhigh)59.1%
- 4Claude Opus 4.8 (xhigh)57.3%
- 5GLM-5.248.1%
- 6Gemini 3.8 Flash47.0%
- 7GPT-5.6 Sol (xhigh)46.0%
- 8GPT-5.5 (xhigh)45.4%
- 9Muse Spark 1.1 (xhigh)42.2%
- 10GPT-5.4 (xhigh)40.8%
- 11Claude Opus 4.740.3%
- 12Claude Fable 5 (xhigh)39.0%
- 13Claude Opus 4.633.3%
- 14GPT-5.3 Chat (xhigh)32.6%
- 15Claude Sonnet 4.631.2%
22 of 22
| # | ||||
|---|---|---|---|---|
| 1 | 63.2% | 70.2 | 2026-08-04 | |
| 2 | 60.0% | 71.6 | 2026-03-02 | |
| 3 | 59.1% | 71.1 | 2026-09-09 | |
| 4 | 57.3% | – | 2026-06-08 | |
| 5 | 48.1% | 52.1 | 2026-06-23 | |
| 6 | 47.0% | – | 2026-09-09 | |
| 7 | 46.0% | 68.2 | 2026-03-02 | |
| 8 | 45.4% | 67.6 | 2026-05-07 | |
| 9 | 42.2% | 61.8 | 2026-07-09 | |
| 10 | 40.8% | 65.4 | 2026-03-09 | |
| 11 | 40.3% | 64.5 | 2026-06-18 | |
| 12 | 39.0% | – | 2026-03-02 | |
| 13 | 33.3% | 63.7 | 2026-02-25 | |
| 14 | 32.6% | – | 2026-02-25 | |
| 15 | 31.2% | 59.0 | 2026-02-25 | |
| 16 | 27.1% | 52.6 | 2026-06-18 | |
| 17 | 24.2% | 65.7 | 2026-04-08 | |
| 18 | 20.5% | 54.7 | 2026-02-25 | |
| 19 | 13.5% | 63.9 | 2026-02-25 | |
| 20 | 13.1% | 53.9 | 2026-02-25 | |
| 21 | 10.3% | 54.2 | 2026-02-25 | |
| 22 | 8.2% | 57.7 | 2026-02-25 |
Cite as: BenchLeader, “SWE Atlas: Codebase QnA leaderboard”, https://www.benchleader.com/benchmarks/scale_sweatlas_qna, data as of 19 Sept 2026.
SWE Atlas: Codebase QnA: questions
- What does SWE Atlas: Codebase QnA measure?
- The model must answer questions about how a real codebase works, requiring navigation and reading across files. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads SWE Atlas: Codebase QnA?
- Claude Opus 5 leads SWE Atlas: Codebase QnA with 63.2% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 60.0%.
- How many models have SWE Atlas: Codebase QnA results?
- 22 model configurations have a SWE Atlas: Codebase QnA result on BenchLeader, all taken from Scale AI SEAL.
- Who runs SWE Atlas: Codebase QnA and how often is it updated?
- SWE Atlas: Codebase QnA is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does SWE Atlas: Codebase QnA count toward the BenchLeader Index?
- No. SWE Atlas: Codebase QnA is shown for reference but left out of the composite index.