BenchLeader

DeepSWE

Software-engineering tasks in an agent harness with reasoning effort recorded.

As of 19 Sept 2026, GPT-6 Astra leads DeepSWE on BenchLeader with 74.1%, ahead of Gemini 3.8 Flash at 73.8%, across 69 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Coding
Index weight
Reference only
Models
69
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

SWE-style repository tasks run in an agent loop, with the reasoning effort noted per run.

How it is scored

Pass@1, as published.

What to keep in mind

Harness-dependent.

69 of 69
#
1GPT-6 AstraxhighOpenAI74.1%71.12026-09-03
2Gemini 3.8 FlashhighGoogle73.8%64.52026-09-02
3Claude Opus 5maxAnthropic73.7%69.92026-07-24
4GPT-6 AstrahighOpenAI73.2%71.92026-09-03
5GPT-6 AstramaxOpenAI73.2%71.82026-09-03
6Claude Opus 5xhighAnthropic73.2%70.22026-07-24
7Claude Opus 5highAnthropic72.8%70.22026-07-24
8GPT-6 AstramediumOpenAI72.8%69.82026-09-03
9GPT-5.6 SolmaxOpenAI72.7%68.82026-07-09
10Gemini 3.8 FlashmediumGoogle71.0%64.92026-09-02
11GPT-5.6 SolxhighOpenAI70.7%68.22026-07-09
12Claude Fable 5xhighAnthropic69.9%2026-06-09
13Claude Fable 5maxAnthropic69.7%65.82026-06-09
14GPT-5.6 TerramaxOpenAI69.6%65.02026-07-09
15GPT-5.6 SolhighOpenAI69.4%68.02026-07-09
16GLM-5.3maxZhipu AIopen ↗69.0%65.82026-08-14
17Claude Opus 5mediumAnthropic68.9%67.32026-07-24
18Claude Fable 5highAnthropic68.6%64.02026-06-09
19Kimi K3maxMoonshot AIopen ↗68.5%67.22026-07-16
20Grok 4.6mediumxAI67.5%66.22026-08-12
21GPT-5.6 LunamaxOpenAI67.2%60.22026-07-09
22GPT-6 AstralowOpenAI67.0%68.22026-09-03
23GPT-5.5xhighOpenAI67.0%67.62026-04-23
24Grok 4.6xhighxAI66.7%64.62026-08-12
25Gemini 3.7 FlashmediumGoogle65.5%64.72026-08-13
26Claude Fable 5mediumAnthropic65.4%2026-06-09
27Gemini 3.7 FlashhighGoogle65.3%64.62026-08-13
28Grok 4.6highxAI65.2%64.52026-08-12
29GPT-5.5highOpenAI64.4%67.02026-04-23
30GLM-5.3-FlashmaxZhipu AIopen ↗63.4%55.32026-08-20
31GPT-5.6 SolmediumOpenAI61.1%66.12026-07-09
32GPT-5.6 TerraxhighOpenAI60.2%64.32026-07-09
33Claude Fable 5lowAnthropic59.6%2026-06-09
34Claude Opus 4.8maxAnthropic59.0%64.22026-05-28
35Claude Opus 5lowAnthropic58.1%63.42026-07-24
36Qwen3 8xhighAlibaba57.5%57.72026-08-02
37GPT-5.6 LunaxhighOpenAI56.9%61.12026-07-09
38Muse Spark 1.2xhighMeta54.9%64.12026-08-05
39Claude Opus 4.8xhighAnthropic54.4%2026-05-28
40GPT-5.5mediumOpenAI54.0%64.62026-04-23
41Claude Sonnet 5maxAnthropic53.9%60.52026-06-30
42GPT-5.6 TerrahighOpenAI53.8%62.82026-07-09
43Gemini 3.7 FlashlowGoogle53.8%62.22026-08-13
44Grok 4.5highxAI53.8%61.02026-07-08
45Muse Spark 1.1Meta53.3%65.22026-07-09
46GPT-5.4xhighOpenAI51.8%65.42026-03-05
47Claude Opus 4.8highAnthropic51.8%62.22026-05-28
48Claude Sonnet 5xhighAnthropic49.7%59.82026-06-30
49Claude Opus 4.8mediumAnthropic48.7%2026-05-28
50Claude Sonnet 5highAnthropic48.2%62.22026-06-30
51Gemini 3.6 FlashhighGoogle46.7%61.72026-07-21
52GPT-5.6 SollowOpenAI45.4%63.72026-07-09
53GPT-5.6 LunahighOpenAI44.3%58.62026-07-09
54GLM-5.2maxZhipu AIopen ↗43.8%63.92026-06-16
55Grok 4.6lowxAI41.6%61.02026-08-12
56Claude Opus 4.8lowAnthropic40.8%2026-05-28
57Claude Sonnet 5mediumAnthropic39.8%2026-06-30
58Gemini 3.5 FlashmediumGoogle37.4%64.22026-05-19
59GLM-5.2highZhipu AIopen ↗36.3%2026-06-16
60Gemini 3.5 FlashhighGoogle36.1%63.62026-05-19

Cite as: BenchLeader, “DeepSWE leaderboard”, https://www.benchleader.com/benchmarks/deepswe, data as of 19 Sept 2026.

DeepSWE: questions

What does DeepSWE measure?
SWE-style repository tasks run in an agent loop, with the reasoning effort noted per run. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads DeepSWE?
GPT-6 Astra leads DeepSWE with 74.1% as of 19 Sept 2026, ahead of Gemini 3.8 Flash at 73.8%.
How many models have DeepSWE results?
69 model configurations have a DeepSWE result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs DeepSWE and how often is it updated?
DeepSWE is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does DeepSWE count toward the BenchLeader Index?
No. DeepSWE is shown for reference but left out of the composite index.