BenchLeader

SWE Atlas: Codebase QnA

Answering questions about a large codebase. Scale AI.

As of 19 Sept 2026, Claude Opus 5 leads SWE Atlas: Codebase QnA on BenchLeader with 63.2%, ahead of Claude Fable 5.1 at 60.0%, across 22 model configurations with a published result.

Published by
Scale AI SEAL
Category
Coding
Index weight
Reference only
Models
22
Data as of
19 Sept 2026

Scale AI SEAL Leaderboards.

What the test looks like

The model must answer questions about how a real codebase works, requiring navigation and reading across files.

How it is scored

Accuracy, published by Scale AI.

What to keep in mind

Measures code comprehension rather than editing.

22 of 22
#
1Claude Opus 5xhighAnthropic63.2%70.22026-08-04
2Claude Fable 5.1xhighAnthropic60.0%71.62026-03-02
3GPT-6 AstraxhighOpenAI59.1%71.12026-09-09
4Claude Opus 4.8xhighAnthropic57.3%2026-06-08
5GLM-5.2Zhipu AIopen ↗48.1%52.12026-06-23
6Gemini 3.8 FlashGoogle47.0%2026-09-09
7GPT-5.6 SolxhighOpenAI46.0%68.22026-03-02
8GPT-5.5xhighOpenAI45.4%67.62026-05-07
9Muse Spark 1.1xhighMeta42.2%61.82026-07-09
10GPT-5.4xhighOpenAI40.8%65.42026-03-09
11Claude Opus 4.7Anthropic40.3%64.52026-06-18
12Claude Fable 5xhighAnthropic39.0%2026-03-02
13Claude Opus 4.6Anthropic33.3%63.72026-02-25
14GPT-5.3 ChatxhighOpenAI32.6%2026-02-25
15Claude Sonnet 4.6Anthropic31.2%59.02026-02-25
16DeepSeek V4 ProDeepSeekopen ↗27.1%52.62026-06-18
17Muse SparkMeta24.2%65.72026-04-08
18GLM-5Zhipu AIopen20.5%54.72026-02-25
19Gemini 3.1 ProGoogle13.5%63.92026-02-25
20Kimi K2.5Moonshot AIopen13.1%53.92026-02-25
21MiniMax-M2.5MiniMaxopen ↗10.3%54.22026-02-25
22Gemini 3 FlashGoogle8.2%57.72026-02-25

Cite as: BenchLeader, “SWE Atlas: Codebase QnA leaderboard”, https://www.benchleader.com/benchmarks/scale_sweatlas_qna, data as of 19 Sept 2026.

SWE Atlas: Codebase QnA: questions

What does SWE Atlas: Codebase QnA measure?
The model must answer questions about how a real codebase works, requiring navigation and reading across files. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads SWE Atlas: Codebase QnA?
Claude Opus 5 leads SWE Atlas: Codebase QnA with 63.2% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 60.0%.
How many models have SWE Atlas: Codebase QnA results?
22 model configurations have a SWE Atlas: Codebase QnA result on BenchLeader, all taken from Scale AI SEAL.
Who runs SWE Atlas: Codebase QnA and how often is it updated?
SWE Atlas: Codebase QnA is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
Does SWE Atlas: Codebase QnA count toward the BenchLeader Index?
No. SWE Atlas: Codebase QnA is shown for reference but left out of the composite index.