BenchLeader

CursorBench

Cursor's internal coding-agent evaluation, with reasoning level and cost.

As of 19 Sept 2026, Claude Fable 5.1 leads CursorBench on BenchLeader with 73.4%, ahead of Claude Fable 5.1 at 72.8%, across 74 model configurations with a published result.

Published by
Cursordata via Epoch AI Benchmarking Hub
Category
Coding
Index weight
Reference only
Models
74
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Coding tasks run in Cursor's agent with the reasoning level and cost per task recorded.

How it is scored

Score, as published by Cursor.

What to keep in mind

Published by a coding-tool vendor; harness is theirs.

74 of 74
#
1Claude Fable 5.1maxAnthropic73.4%69.72026-09-01
2Claude Fable 5.1xhighAnthropic72.8%71.62026-09-01
3Grok 4.6xhighxAI70.8%64.62026-08-12
4Claude Fable 5maxAnthropic70.5%65.82026-06-09
5Claude Opus 5maxAnthropic70.0%69.92026-07-24
6Grok 4.6highxAI69.9%64.52026-08-12
7Claude Fable 5.1highAnthropic69.4%72.02026-09-01
8Claude Opus 5xhighAnthropic69.3%70.22026-07-24
9Gemini 3.8 FlashhighGoogle69.2%64.52026-09-02
10Claude Fable 5xhighAnthropic68.4%2026-06-09
11Claude Fable 5.1mediumAnthropic68.0%69.62026-09-01
12GPT-5.6 SolmaxOpenAI67.2%68.82026-07-09
13Grok 4.6mediumxAI67.1%66.22026-08-12
14Gemini 3.8 FlashmediumGoogle67.0%64.92026-09-02
15Claude Opus 5highAnthropic66.7%70.22026-07-24
16Claude Fable 5highAnthropic66.5%64.02026-06-09
17Claude Fable 5.1lowAnthropic66.2%67.52026-09-01
18Claude Fable 5mediumAnthropic65.2%2026-06-09
19GPT-5.6 TerramaxOpenAI64.9%65.02026-07-09
20Claude Opus 4.7maxAnthropic64.8%64.32026-04-16
21GPT-5.6 SolxhighOpenAI64.5%68.22026-07-09
22Claude Opus 5mediumAnthropic64.3%67.32026-07-24
23GPT-5.6 SolhighOpenAI63.5%68.02026-07-09
24Claude Opus 5lowAnthropic62.8%63.42026-07-24
25Claude Opus 4.8maxAnthropic62.3%64.22026-05-28
26Claude Fable 5lowAnthropic62.1%2026-06-09
27Gemini 3.7 FlashhighGoogle61.6%64.62026-08-13
28Claude Opus 4.7xhighAnthropic61.6%57.32026-04-16
29Claude Sonnet 5maxAnthropic61.5%60.52026-06-30
30GPT-5.6 LunamaxOpenAI61.1%60.22026-07-09
31Grok 4.6lowxAI61.0%61.02026-08-12
32Kimi K3maxMoonshot AIopen ↗60.8%67.22026-07-16
33GPT-5.6 SolmediumOpenAI60.0%66.12026-07-09
34Kimi K3highMoonshot AIopen ↗59.7%2026-07-16
35Claude Opus 4.7highAnthropic59.4%62.42026-04-16
36Claude Opus 4.8xhighAnthropic59.4%2026-05-28
37GPT-5.6 TerraxhighOpenAI59.2%64.32026-07-09
38Gemini 3.7 FlashmediumGoogle59.0%64.72026-08-13
39Claude Sonnet 5xhighAnthropic58.7%59.82026-06-30
40GPT-5.5xhighOpenAI58.4%67.62026-04-23
41GPT-5.5highOpenAI58.4%67.02026-04-23
42Claude Opus 4.8highAnthropic58.0%62.22026-05-28
43GPT-5.6 LunaxhighOpenAI57.7%61.12026-07-09
44Claude Sonnet 5highAnthropic56.9%62.22026-06-30
45GPT-5.6 LunahighOpenAI56.8%58.62026-07-09
46Claude Opus 4.8mediumAnthropic56.1%2026-05-28
47Composer 2.5Cursor56.1%2026-05-18
48GLM-5.2maxZhipu AIopen ↗55.0%63.92026-06-16
49GPT-5.6 TerrahighOpenAI54.2%62.82026-07-09
50GPT-5.5mediumOpenAI53.8%64.62026-04-23
51Gemini 3.7 FlashlowGoogle53.8%62.22026-08-13
52Gemini 3.6 FlashhighGoogle53.5%61.72026-07-21
53Claude Opus 4.8lowAnthropic53.1%2026-05-28
54Claude Opus 4.7mediumAnthropic52.7%2026-04-16
55GPT-5.6 SollowOpenAI52.6%63.72026-07-09
56Claude Sonnet 5mediumAnthropic52.4%2026-06-30
57GLM-5.2highZhipu AIopen ↗51.5%2026-06-16
58Gemini 3.6 FlashmediumGoogle51.2%2026-07-21
59Kimi K3lowMoonshot AIopen ↗50.5%59.62026-07-16
60GPT-5.6 TerramediumOpenAI50.3%58.42026-07-09

Cite as: BenchLeader, “CursorBench leaderboard”, https://www.benchleader.com/benchmarks/cursorbench, data as of 19 Sept 2026.

CursorBench: questions

What does CursorBench measure?
Coding tasks run in Cursor's agent with the reasoning level and cost per task recorded. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads CursorBench?
Claude Fable 5.1 leads CursorBench with 73.4% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 72.8%.
How many models have CursorBench results?
74 model configurations have a CursorBench result on BenchLeader, all taken from Cursor via Epoch AI Benchmarking Hub.
Who runs CursorBench and how often is it updated?
CursorBench is published by Cursor. BenchLeader re-reads the published results every morning and records the date each result was published.
Does CursorBench count toward the BenchLeader Index?
No. CursorBench is shown for reference but left out of the composite index.