CadEval
Generating CAD models from specifications.
As of 19 Sept 2026, o3 leads CadEval on BenchLeader with 74.0%, ahead of Gemini 2.5 Pro at 64.0%, across 15 model configurations with a published result.
- Published by
- Epoch AI Benchmarking Hub
- Category
- Coding
- Index weight
- Reference only
- Models
- 15
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
The model writes CAD code to produce a part matching a specification, checked geometrically.
How it is scored
Overall pass rate, as published.
What to keep in mind
Small set; a niche but checkable skill.
- 1o3 (medium)74.0%
- 2Gemini 2.5 Pro64.0%
- 3o4-mini (medium)62.0%
- 4o1 (medium)56.0%
- 5Claude 3.7 Sonnet54.0%
- 6o3-mini (medium)54.0%
- 7Claude 3.5 Sonnet48.0%
- 8GPT-4.142.0%
- 9Gemini 1.5 Pro 00234.0%
- 10Claude 3.5 Haiku32.0%
- 11Gemini 2.0 Flash30.0%
- 12GPT-4o26.0%
- 13ml-elephant20.0%
- 14GPT-4.1 mini16.0%
- 15Claude 3 Haiku12.0%
15 of 15
| # | ||||
|---|---|---|---|---|
| 1 | 74.0% | 53.7 | 2025-04-16 | |
| 2 | 64.0% | 54.5 | 2025-03-31 | |
| 3 | 62.0% | 52.4 | 2025-04-16 | |
| 4 | 56.0% | 48.4 | 2024-12-17 | |
| 5 | 54.0% | 50.2 | 2025-02-24 | |
| 6 | 54.0% | 45.6 | 2025-01-31 | |
| 7 | 48.0% | 46.6 | 2024-10-22 | |
| 8 | 42.0% | 49.0 | 2025-04-14 | |
| 9 | 34.0% | 44.9 | 2024-09-24 | |
| 10 | 32.0% | 40.3 | 2024-10-22 | |
| 11 | 30.0% | 43.7 | 2025-02-05 | |
| 12 | 26.0% | 44.3 | 2024-08-06 | |
| 13 | ml-elephantUnknown | 20.0% | – | – |
| 14 | 16.0% | 45.6 | 2025-04-14 | |
| 15 | 12.0% | 37.6 | 2024-03-07 |
Cite as: BenchLeader, “CadEval leaderboard”, https://www.benchleader.com/benchmarks/cad_eval, data as of 19 Sept 2026.
CadEval: questions
- What does CadEval measure?
- The model writes CAD code to produce a part matching a specification, checked geometrically. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads CadEval?
- o3 leads CadEval with 74.0% as of 19 Sept 2026, ahead of Gemini 2.5 Pro at 64.0%.
- How many models have CadEval results?
- 15 model configurations have a CadEval result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
- Who runs CadEval and how often is it updated?
- CadEval is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does CadEval count toward the BenchLeader Index?
- No. CadEval is shown for reference but left out of the composite index.