BenchLeader

DrugDiscoveryBench

Drug-discovery reasoning tasks graded by practitioners. Scale AI.

As of 19 Sept 2026, GPT-6 Astra leads DrugDiscoveryBench on BenchLeader with 68.7%, ahead of Muse Spark 1.3 at 62.2%, across 18 model configurations with a published result.

Published by
Scale AI SEAL
Category
Knowledge
Index weight
Reference only
Models
18
Data as of
19 Sept 2026

Scale AI SEAL Leaderboards.

What the test looks like

Questions and tasks from the drug-discovery pipeline, from target selection to trial design, graded by domain scientists.

How it is scored

Rubric score, published by Scale AI.

What to keep in mind

Expert-graded and narrow; small model set.

18 of 18
#
1GPT-6 AstramaxOpenAI68.7%71.82026-09-14
2Muse Spark 1.3highMeta62.2%2026-09-14
3GPT-5.6 SolmaxOpenAI56.5%68.82026-09-11
4Claude Opus 5maxAnthropic53.7%69.92026-09-11
5GPT-5.5xhighOpenAI51.6%67.62026-06-30
6Claude Sonnet 5maxAnthropic50.0%60.52026-07-14
7Gemini 3.5 FlashhighGoogle50.0%63.62026-06-30
8Claude Opus 4.8maxAnthropic46.8%64.22026-06-30
9Claude Sonnet 5.0maxAnthropic44.7%2026-07-14
10Gemini 3.1 ProhighGoogle41.9%60.72026-06-30
11GLM-5.2xhighZhipu AIopen ↗36.2%2026-06-30
12Kimi K2.7 CodexhighMoonshot AIopen ↗35.3%2026-06-30
13DeepSeek V4 ProxhighDeepSeekopen ↗31.7%2026-06-30
14Claude Sonnet 4.6maxAnthropic31.3%57.62026-06-30
15GPT-5.2xhighOpenAI29.3%62.02026-06-30
16Qwen3 7xhighAlibaba29.3%2026-06-30
17Claude Opus 4.6maxAnthropic27.7%59.22026-06-30
18MiniMax-M3xhighMiniMaxopen ↗22.8%2026-06-30

Cite as: BenchLeader, “DrugDiscoveryBench leaderboard”, https://www.benchleader.com/benchmarks/scale_drugdiscovery, data as of 19 Sept 2026.

DrugDiscoveryBench: questions

What does DrugDiscoveryBench measure?
Questions and tasks from the drug-discovery pipeline, from target selection to trial design, graded by domain scientists. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads DrugDiscoveryBench?
GPT-6 Astra leads DrugDiscoveryBench with 68.7% as of 19 Sept 2026, ahead of Muse Spark 1.3 at 62.2%.
How many models have DrugDiscoveryBench results?
18 model configurations have a DrugDiscoveryBench result on BenchLeader, all taken from Scale AI SEAL.
Who runs DrugDiscoveryBench and how often is it updated?
DrugDiscoveryBench is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
Does DrugDiscoveryBench count toward the BenchLeader Index?
No. DrugDiscoveryBench is shown for reference but left out of the composite index.