SWE-Bench Pro (private)
The held-out commercial-repository set of SWE-Bench Pro. Scale AI.
As of 19 Sept 2026, Muse Spark 1.1 leads SWE-Bench Pro (private) on BenchLeader with 51.5%, ahead of Claude Opus 4.6 at 47.1%, across 14 model configurations with a published result.
- Published by
- Scale AI SEAL
- Category
- Coding
- Index weight
- Reference only
- Models
- 14
- Data as of
- 19 Sept 2026
Scale AI SEAL Leaderboards.
What the test looks like
The same task format as SWE-Bench Pro on a held-out set of commercial repositories only Scale can run.
How it is scored
Percent of issues resolved, published by Scale AI.
What to keep in mind
Contamination-proof but not reproducible outside Scale.
- 1Muse Spark 1.151.5%
- 2Claude Opus 4.6 (thinking)47.1%
- 3Muse Spark44.7%
- 4GPT-5.4 (xhigh)43.4%
- 5Gemini 3.1 Pro (thinking)32.2%
- 6GPT-5.2-Codex27.7%
- 7GPT-5.223.8%
- 8Claude Opus 4.523.4%
- 9Gemini 3 Pro17.9%
- 10Claude Opus 4.117.8%
- 11GPT-514.9%
- 12Gemini 2.5 Pro10.1%
- 13Claude Sonnet 49.1%
- 14GPT-4o3.6%
14 of 14
| # | ||||
|---|---|---|---|---|
| 1 | 51.5% | 65.2 | 2026-07-09 | |
| 2 | 47.1% | 62.5 | 2026-04-08 | |
| 3 | 44.7% | 65.7 | 2026-04-08 | |
| 4 | 43.4% | 65.4 | 2026-04-08 | |
| 5 | 32.2% | – | 2026-04-08 | |
| 6 | 27.7% | 54.0 | 2026-01-12 | |
| 7 | 23.8% | 58.4 | 2026-01-12 | |
| 8 | 23.4% | 58.2 | 2026-01-12 | |
| 9 | 17.9% | 61.1 | 2025-09-19 | |
| 10 | 17.8% | 51.9 | 2025-09-19 | |
| 11 | 14.9% | 57.2 | 2025-09-19 | |
| 12 | 10.1% | 54.5 | 2025-09-19 | |
| 13 | 9.1% | 49.7 | 2025-09-19 | |
| 14 | 3.6% | 44.3 | 2025-09-19 |
Cite as: BenchLeader, “SWE-Bench Pro (private) leaderboard”, https://www.benchleader.com/benchmarks/scale_swe_bench_pro_private, data as of 19 Sept 2026.
SWE-Bench Pro (private): questions
- What does SWE-Bench Pro (private) measure?
- The same task format as SWE-Bench Pro on a held-out set of commercial repositories only Scale can run. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads SWE-Bench Pro (private)?
- Muse Spark 1.1 leads SWE-Bench Pro (private) with 51.5% as of 19 Sept 2026, ahead of Claude Opus 4.6 at 47.1%.
- How many models have SWE-Bench Pro (private) results?
- 14 model configurations have a SWE-Bench Pro (private) result on BenchLeader, all taken from Scale AI SEAL.
- Who runs SWE-Bench Pro (private) and how often is it updated?
- SWE-Bench Pro (private) is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does SWE-Bench Pro (private) count toward the BenchLeader Index?
- No. SWE-Bench Pro (private) is shown for reference but left out of the composite index.