CursorBench
Cursor's internal coding-agent evaluation, with reasoning level and cost.
As of 19 Sept 2026, Claude Fable 5.1 leads CursorBench on BenchLeader with 73.4%, ahead of Claude Fable 5.1 at 72.8%, across 74 model configurations with a published result.
- Published by
- Cursordata via Epoch AI Benchmarking Hub
- Category
- Coding
- Index weight
- Reference only
- Models
- 74
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Coding tasks run in Cursor's agent with the reasoning level and cost per task recorded.
How it is scored
Score, as published by Cursor.
What to keep in mind
Published by a coding-tool vendor; harness is theirs.
- 1Claude Fable 5.1 (max)73.4%
- 2Claude Fable 5.1 (xhigh)72.8%
- 3Grok 4.6 (xhigh)70.8%
- 4Claude Fable 5 (max)70.5%
- 5Claude Opus 5 (max)70.0%
- 6Grok 4.6 (high)69.9%
- 7Claude Fable 5.1 (high)69.4%
- 8Claude Opus 5 (xhigh)69.3%
- 9Gemini 3.8 Flash (high)69.2%
- 10Claude Fable 5 (xhigh)68.4%
- 11Claude Fable 5.1 (medium)68.0%
- 12GPT-5.6 Sol (max)67.2%
- 13Grok 4.6 (medium)67.1%
- 14Gemini 3.8 Flash (medium)67.0%
- 15Claude Opus 5 (high)66.7%
74 of 74
| # | ||||
|---|---|---|---|---|
| 1 | 73.4% | 69.7 | 2026-09-01 | |
| 2 | 72.8% | 71.6 | 2026-09-01 | |
| 3 | 70.8% | 64.6 | 2026-08-12 | |
| 4 | 70.5% | 65.8 | 2026-06-09 | |
| 5 | 70.0% | 69.9 | 2026-07-24 | |
| 6 | 69.9% | 64.5 | 2026-08-12 | |
| 7 | 69.4% | 72.0 | 2026-09-01 | |
| 8 | 69.3% | 70.2 | 2026-07-24 | |
| 9 | 69.2% | 64.5 | 2026-09-02 | |
| 10 | 68.4% | – | 2026-06-09 | |
| 11 | 68.0% | 69.6 | 2026-09-01 | |
| 12 | 67.2% | 68.8 | 2026-07-09 | |
| 13 | 67.1% | 66.2 | 2026-08-12 | |
| 14 | 67.0% | 64.9 | 2026-09-02 | |
| 15 | 66.7% | 70.2 | 2026-07-24 | |
| 16 | 66.5% | 64.0 | 2026-06-09 | |
| 17 | 66.2% | 67.5 | 2026-09-01 | |
| 18 | 65.2% | – | 2026-06-09 | |
| 19 | 64.9% | 65.0 | 2026-07-09 | |
| 20 | 64.8% | 64.3 | 2026-04-16 | |
| 21 | 64.5% | 68.2 | 2026-07-09 | |
| 22 | 64.3% | 67.3 | 2026-07-24 | |
| 23 | 63.5% | 68.0 | 2026-07-09 | |
| 24 | 62.8% | 63.4 | 2026-07-24 | |
| 25 | 62.3% | 64.2 | 2026-05-28 | |
| 26 | 62.1% | – | 2026-06-09 | |
| 27 | 61.6% | 64.6 | 2026-08-13 | |
| 28 | 61.6% | 57.3 | 2026-04-16 | |
| 29 | 61.5% | 60.5 | 2026-06-30 | |
| 30 | 61.1% | 60.2 | 2026-07-09 | |
| 31 | 61.0% | 61.0 | 2026-08-12 | |
| 32 | 60.8% | 67.2 | 2026-07-16 | |
| 33 | 60.0% | 66.1 | 2026-07-09 | |
| 34 | 59.7% | – | 2026-07-16 | |
| 35 | 59.4% | 62.4 | 2026-04-16 | |
| 36 | 59.4% | – | 2026-05-28 | |
| 37 | 59.2% | 64.3 | 2026-07-09 | |
| 38 | 59.0% | 64.7 | 2026-08-13 | |
| 39 | 58.7% | 59.8 | 2026-06-30 | |
| 40 | 58.4% | 67.6 | 2026-04-23 | |
| 41 | 58.4% | 67.0 | 2026-04-23 | |
| 42 | 58.0% | 62.2 | 2026-05-28 | |
| 43 | 57.7% | 61.1 | 2026-07-09 | |
| 44 | 56.9% | 62.2 | 2026-06-30 | |
| 45 | 56.8% | 58.6 | 2026-07-09 | |
| 46 | 56.1% | – | 2026-05-28 | |
| 47 | 56.1% | – | 2026-05-18 | |
| 48 | 55.0% | 63.9 | 2026-06-16 | |
| 49 | 54.2% | 62.8 | 2026-07-09 | |
| 50 | 53.8% | 64.6 | 2026-04-23 | |
| 51 | 53.8% | 62.2 | 2026-08-13 | |
| 52 | 53.5% | 61.7 | 2026-07-21 | |
| 53 | 53.1% | – | 2026-05-28 | |
| 54 | 52.7% | – | 2026-04-16 | |
| 55 | 52.6% | 63.7 | 2026-07-09 | |
| 56 | 52.4% | – | 2026-06-30 | |
| 57 | 51.5% | – | 2026-06-16 | |
| 58 | 51.2% | – | 2026-07-21 | |
| 59 | 50.5% | 59.6 | 2026-07-16 | |
| 60 | 50.3% | 58.4 | 2026-07-09 |
Cite as: BenchLeader, “CursorBench leaderboard”, https://www.benchleader.com/benchmarks/cursorbench, data as of 19 Sept 2026.
CursorBench: questions
- What does CursorBench measure?
- Coding tasks run in Cursor's agent with the reasoning level and cost per task recorded. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads CursorBench?
- Claude Fable 5.1 leads CursorBench with 73.4% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 72.8%.
- How many models have CursorBench results?
- 74 model configurations have a CursorBench result on BenchLeader, all taken from Cursor via Epoch AI Benchmarking Hub.
- Who runs CursorBench and how often is it updated?
- CursorBench is published by Cursor. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does CursorBench count toward the BenchLeader Index?
- No. CursorBench is shown for reference but left out of the composite index.