DeepSWE
Software-engineering tasks in an agent harness with reasoning effort recorded.
As of 19 Sept 2026, GPT-6 Astra leads DeepSWE on BenchLeader with 74.1%, ahead of Gemini 3.8 Flash at 73.8%, across 69 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Coding
- Index weight
- Reference only
- Models
- 69
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
SWE-style repository tasks run in an agent loop, with the reasoning effort noted per run.
How it is scored
Pass@1, as published.
What to keep in mind
Harness-dependent.
- 1GPT-6 Astra (xhigh)74.1%
- 2Gemini 3.8 Flash (high)73.8%
- 3Claude Opus 5 (max)73.7%
- 4GPT-6 Astra (high)73.2%
- 5GPT-6 Astra (max)73.2%
- 6Claude Opus 5 (xhigh)73.2%
- 7Claude Opus 5 (high)72.8%
- 8GPT-6 Astra (medium)72.8%
- 9GPT-5.6 Sol (max)72.7%
- 10Gemini 3.8 Flash (medium)71.0%
- 11GPT-5.6 Sol (xhigh)70.7%
- 12Claude Fable 5 (xhigh)69.9%
- 13Claude Fable 5 (max)69.7%
- 14GPT-5.6 Terra (max)69.6%
- 15GPT-5.6 Sol (high)69.4%
69 of 69
| # | ||||
|---|---|---|---|---|
| 1 | 74.1% | 71.1 | 2026-09-03 | |
| 2 | 73.8% | 64.5 | 2026-09-02 | |
| 3 | 73.7% | 69.9 | 2026-07-24 | |
| 4 | 73.2% | 71.9 | 2026-09-03 | |
| 5 | 73.2% | 71.8 | 2026-09-03 | |
| 6 | 73.2% | 70.2 | 2026-07-24 | |
| 7 | 72.8% | 70.2 | 2026-07-24 | |
| 8 | 72.8% | 69.8 | 2026-09-03 | |
| 9 | 72.7% | 68.8 | 2026-07-09 | |
| 10 | 71.0% | 64.9 | 2026-09-02 | |
| 11 | 70.7% | 68.2 | 2026-07-09 | |
| 12 | 69.9% | – | 2026-06-09 | |
| 13 | 69.7% | 65.8 | 2026-06-09 | |
| 14 | 69.6% | 65.0 | 2026-07-09 | |
| 15 | 69.4% | 68.0 | 2026-07-09 | |
| 16 | 69.0% | 65.8 | 2026-08-14 | |
| 17 | 68.9% | 67.3 | 2026-07-24 | |
| 18 | 68.6% | 64.0 | 2026-06-09 | |
| 19 | 68.5% | 67.2 | 2026-07-16 | |
| 20 | 67.5% | 66.2 | 2026-08-12 | |
| 21 | 67.2% | 60.2 | 2026-07-09 | |
| 22 | 67.0% | 68.2 | 2026-09-03 | |
| 23 | 67.0% | 67.6 | 2026-04-23 | |
| 24 | 66.7% | 64.6 | 2026-08-12 | |
| 25 | 65.5% | 64.7 | 2026-08-13 | |
| 26 | 65.4% | – | 2026-06-09 | |
| 27 | 65.3% | 64.6 | 2026-08-13 | |
| 28 | 65.2% | 64.5 | 2026-08-12 | |
| 29 | 64.4% | 67.0 | 2026-04-23 | |
| 30 | 63.4% | 55.3 | 2026-08-20 | |
| 31 | 61.1% | 66.1 | 2026-07-09 | |
| 32 | 60.2% | 64.3 | 2026-07-09 | |
| 33 | 59.6% | – | 2026-06-09 | |
| 34 | 59.0% | 64.2 | 2026-05-28 | |
| 35 | 58.1% | 63.4 | 2026-07-24 | |
| 36 | 57.5% | 57.7 | 2026-08-02 | |
| 37 | 56.9% | 61.1 | 2026-07-09 | |
| 38 | 54.9% | 64.1 | 2026-08-05 | |
| 39 | 54.4% | – | 2026-05-28 | |
| 40 | 54.0% | 64.6 | 2026-04-23 | |
| 41 | 53.9% | 60.5 | 2026-06-30 | |
| 42 | 53.8% | 62.8 | 2026-07-09 | |
| 43 | 53.8% | 62.2 | 2026-08-13 | |
| 44 | 53.8% | 61.0 | 2026-07-08 | |
| 45 | 53.3% | 65.2 | 2026-07-09 | |
| 46 | 51.8% | 65.4 | 2026-03-05 | |
| 47 | 51.8% | 62.2 | 2026-05-28 | |
| 48 | 49.7% | 59.8 | 2026-06-30 | |
| 49 | 48.7% | – | 2026-05-28 | |
| 50 | 48.2% | 62.2 | 2026-06-30 | |
| 51 | 46.7% | 61.7 | 2026-07-21 | |
| 52 | 45.4% | 63.7 | 2026-07-09 | |
| 53 | 44.3% | 58.6 | 2026-07-09 | |
| 54 | 43.8% | 63.9 | 2026-06-16 | |
| 55 | 41.6% | 61.0 | 2026-08-12 | |
| 56 | 40.8% | – | 2026-05-28 | |
| 57 | 39.8% | – | 2026-06-30 | |
| 58 | 37.4% | 64.2 | 2026-05-19 | |
| 59 | 36.3% | – | 2026-06-16 | |
| 60 | 36.1% | 63.6 | 2026-05-19 |
Cite as: BenchLeader, “DeepSWE leaderboard”, https://www.benchleader.com/benchmarks/deepswe, data as of 19 Sept 2026.
DeepSWE: questions
- What does DeepSWE measure?
- SWE-style repository tasks run in an agent loop, with the reasoning effort noted per run. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads DeepSWE?
- GPT-6 Astra leads DeepSWE with 74.1% as of 19 Sept 2026, ahead of Gemini 3.8 Flash at 73.8%.
- How many models have DeepSWE results?
- 69 model configurations have a DeepSWE result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs DeepSWE and how often is it updated?
- DeepSWE is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does DeepSWE count toward the BenchLeader Index?
- No. DeepSWE is shown for reference but left out of the composite index.