BenchLeader

SWE-Bench Pro (private)

The held-out commercial-repository set of SWE-Bench Pro. Scale AI.

As of 19 Sept 2026, Muse Spark 1.1 leads SWE-Bench Pro (private) on BenchLeader with 51.5%, ahead of Claude Opus 4.6 at 47.1%, across 14 model configurations with a published result.

Published by
Scale AI SEAL
Category
Coding
Index weight
Reference only
Models
14
Data as of
19 Sept 2026

Scale AI SEAL Leaderboards.

What the test looks like

The same task format as SWE-Bench Pro on a held-out set of commercial repositories only Scale can run.

How it is scored

Percent of issues resolved, published by Scale AI.

What to keep in mind

Contamination-proof but not reproducible outside Scale.

14 of 14
#
1Muse Spark 1.1Meta51.5%65.22026-07-09
2Claude Opus 4.6thinkingAnthropic47.1%62.52026-04-08
3Muse SparkMeta44.7%65.72026-04-08
4GPT-5.4xhighOpenAI43.4%65.42026-04-08
5Gemini 3.1 ProthinkingGoogle32.2%2026-04-08
6GPT-5.2-CodexOpenAI27.7%54.02026-01-12
7GPT-5.2OpenAI23.8%58.42026-01-12
8Claude Opus 4.5Anthropic23.4%58.22026-01-12
9Gemini 3 ProGoogle17.9%61.12025-09-19
10Claude Opus 4.1Anthropic17.8%51.92025-09-19
11GPT-5OpenAI14.9%57.22025-09-19
12Gemini 2.5 ProGoogle10.1%54.52025-09-19
13Claude Sonnet 4Anthropic9.1%49.72025-09-19
14GPT-4oOpenAI3.6%44.32025-09-19

Cite as: BenchLeader, “SWE-Bench Pro (private) leaderboard”, https://www.benchleader.com/benchmarks/scale_swe_bench_pro_private, data as of 19 Sept 2026.

SWE-Bench Pro (private): questions

What does SWE-Bench Pro (private) measure?
The same task format as SWE-Bench Pro on a held-out set of commercial repositories only Scale can run. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads SWE-Bench Pro (private)?
Muse Spark 1.1 leads SWE-Bench Pro (private) with 51.5% as of 19 Sept 2026, ahead of Claude Opus 4.6 at 47.1%.
How many models have SWE-Bench Pro (private) results?
14 model configurations have a SWE-Bench Pro (private) result on BenchLeader, all taken from Scale AI SEAL.
Who runs SWE-Bench Pro (private) and how often is it updated?
SWE-Bench Pro (private) is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
Does SWE-Bench Pro (private) count toward the BenchLeader Index?
No. SWE-Bench Pro (private) is shown for reference but left out of the composite index.