FrontierSWE
Frontier-difficulty software-engineering tasks across implementation, performance and research.
As of 19 Sept 2026, Claude Fable 5.1 leads FrontierSWE on BenchLeader with 56.3%, ahead of GPT-5.6 Sol at 32.2%, across 8 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Coding
- Index weight
- Reference only
- Models
- 8
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Long software tasks split into implementation, performance and research work, run in an agent harness.
How it is scored
Score, as published.
What to keep in mind
Very new; a dozen models.
- 1Claude Fable 5.1 (max)56.3%
- 2GPT-5.6 Sol (max)32.2%
- 3GLM-5.3 (max)30.2%
- 4Kimi K3 (max)25.9%
- 5Grok 4.6 (xhigh)25.3%
- 6Gemini 3.7 Flash (high)20.3%
- 7Qwen3 8 (xhigh)15.8%
- 8Muse Spark 1.2 (xhigh)12.0%
8 of 8
| # | ||||
|---|---|---|---|---|
| 1 | 56.3% | 69.7 | 2026-09-01 | |
| 2 | 32.2% | 68.8 | 2026-07-09 | |
| 3 | 30.2% | 65.8 | 2026-08-14 | |
| 4 | 25.9% | 67.2 | 2026-07-16 | |
| 5 | 25.3% | 64.6 | 2026-08-12 | |
| 6 | 20.3% | 64.6 | 2026-08-13 | |
| 7 | 15.8% | 57.7 | 2026-08-02 | |
| 8 | 12.0% | 64.1 | 2026-08-05 |
Cite as: BenchLeader, “FrontierSWE leaderboard”, https://www.benchleader.com/benchmarks/frontierswe, data as of 19 Sept 2026.
FrontierSWE: questions
- What does FrontierSWE measure?
- Long software tasks split into implementation, performance and research work, run in an agent harness. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads FrontierSWE?
- Claude Fable 5.1 leads FrontierSWE with 56.3% as of 19 Sept 2026, ahead of GPT-5.6 Sol at 32.2%.
- How many models have FrontierSWE results?
- 8 model configurations have a FrontierSWE result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs FrontierSWE and how often is it updated?
- FrontierSWE is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does FrontierSWE count toward the BenchLeader Index?
- No. FrontierSWE is shown for reference but left out of the composite index.