DrugDiscoveryBench
Drug-discovery reasoning tasks graded by practitioners. Scale AI.
As of 19 Sept 2026, GPT-6 Astra leads DrugDiscoveryBench on BenchLeader with 68.7%, ahead of Muse Spark 1.3 at 62.2%, across 18 model configurations with a published result.
- Published by
- Scale AI SEAL
- Category
- Knowledge
- Index weight
- Reference only
- Models
- 18
- Data as of
- 19 Sept 2026
Scale AI SEAL Leaderboards.
What the test looks like
Questions and tasks from the drug-discovery pipeline, from target selection to trial design, graded by domain scientists.
How it is scored
Rubric score, published by Scale AI.
What to keep in mind
Expert-graded and narrow; small model set.
- 1GPT-6 Astra (max)68.7%
- 2Muse Spark 1.3 (high)62.2%
- 3GPT-5.6 Sol (max)56.5%
- 4Claude Opus 5 (max)53.7%
- 5GPT-5.5 (xhigh)51.6%
- 6Claude Sonnet 5 (max)50.0%
- 7Gemini 3.5 Flash (high)50.0%
- 8Claude Opus 4.8 (max)46.8%
- 9Claude Sonnet 5.0 (max)44.7%
- 10Gemini 3.1 Pro (high)41.9%
- 11GLM-5.2 (xhigh)36.2%
- 12Kimi K2.7 Code (xhigh)35.3%
- 13DeepSeek V4 Pro (xhigh)31.7%
- 14Claude Sonnet 4.6 (max)31.3%
- 15GPT-5.2 (xhigh)29.3%
18 of 18
| # | ||||
|---|---|---|---|---|
| 1 | 68.7% | 71.8 | 2026-09-14 | |
| 2 | 62.2% | – | 2026-09-14 | |
| 3 | 56.5% | 68.8 | 2026-09-11 | |
| 4 | 53.7% | 69.9 | 2026-09-11 | |
| 5 | 51.6% | 67.6 | 2026-06-30 | |
| 6 | 50.0% | 60.5 | 2026-07-14 | |
| 7 | 50.0% | 63.6 | 2026-06-30 | |
| 8 | 46.8% | 64.2 | 2026-06-30 | |
| 9 | 44.7% | – | 2026-07-14 | |
| 10 | 41.9% | 60.7 | 2026-06-30 | |
| 11 | 36.2% | – | 2026-06-30 | |
| 12 | 35.3% | – | 2026-06-30 | |
| 13 | 31.7% | – | 2026-06-30 | |
| 14 | 31.3% | 57.6 | 2026-06-30 | |
| 15 | 29.3% | 62.0 | 2026-06-30 | |
| 16 | 29.3% | – | 2026-06-30 | |
| 17 | 27.7% | 59.2 | 2026-06-30 | |
| 18 | 22.8% | – | 2026-06-30 |
Cite as: BenchLeader, “DrugDiscoveryBench leaderboard”, https://www.benchleader.com/benchmarks/scale_drugdiscovery, data as of 19 Sept 2026.
DrugDiscoveryBench: questions
- What does DrugDiscoveryBench measure?
- Questions and tasks from the drug-discovery pipeline, from target selection to trial design, graded by domain scientists. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads DrugDiscoveryBench?
- GPT-6 Astra leads DrugDiscoveryBench with 68.7% as of 19 Sept 2026, ahead of Muse Spark 1.3 at 62.2%.
- How many models have DrugDiscoveryBench results?
- 18 model configurations have a DrugDiscoveryBench result on BenchLeader, all taken from Scale AI SEAL.
- Who runs DrugDiscoveryBench and how often is it updated?
- DrugDiscoveryBench is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does DrugDiscoveryBench count toward the BenchLeader Index?
- No. DrugDiscoveryBench is shown for reference but left out of the composite index.