FORTRESS
Robustness to harmful requests without over-refusing benign ones. Scale AI.
As of 19 Sept 2026, DeepSeek R1 leads FORTRESS on BenchLeader with 74.4%, ahead of GLM-4.5-Air at 63.2%, across 62 model configurations with a published result.
- Published by
- Scale AI SEAL
- Category
- Safety & honesty
- Index weight
- Reference only
- Models
- 62
- Data as of
- 19 Sept 2026
Scale AI SEAL Leaderboards.
What the test looks like
Paired adversarial and benign prompts across risk areas; a model must refuse the harmful ones and answer the benign ones.
How it is scored
Combined score for harmful-request refusal and benign-request compliance, published by Scale AI.
What to keep in mind
A safety measure, not a capability one; it is not part of the BenchLeader Index.
- 1DeepSeek R174.4%
- 2GLM-4.5-Air63.2%
- 3Gemini 2.5 Pro61.7%
- 4Qwen3 235B A22B61.4%
- 5GLM-4.559.6%
- 6GPT-4.1 mini59.2%
- 7Qwen2.5 72B56.4%
- 8Mixtral 8x22B Instruct56.1%
- 9Kimi K255.5%
- 10Gemini 1.5 Pro53.9%
- 11GPT-4.153.0%
- 12Gemini 3.1 Flash Lite51.0%
- 13Gemini 1.5 Flash50.6%
- 14GPT-4o mini48.1%
- 15GPT-4o47.2%
62 of 62
| # | ||||
|---|---|---|---|---|
| 1 | 74.4% | 48.7 | 2025-04-10 | |
| 2 | 63.2% | 47.2 | 2025-08-13 | |
| 3 | 61.7% | 54.5 | 2025-04-14 | |
| 4 | 61.4% | 49.6 | 2025-07-24 | |
| 5 | 59.6% | 52.1 | 2025-08-13 | |
| 6 | 59.2% | 45.6 | 2025-04-10 | |
| 7 | 56.4% | 42.7 | 2025-04-10 | |
| 8 | 56.1% | 37.7 | 2025-05-23 | |
| 9 | 55.5% | 51.0 | 2025-07-23 | |
| 10 | 53.9% | 42.9 | 2025-04-10 | |
| 11 | 53.0% | 49.0 | 2025-05-23 | |
| 12 | 51.0% | 54.5 | 2026-03-13 | |
| 13 | 50.6% | – | 2025-04-10 | |
| 14 | 48.1% | 36.2 | 2025-04-10 | |
| 15 | 47.2% | 44.3 | 2025-04-10 | |
| 16 | 44.8% | 41.4 | 2025-04-10 | |
| 17 | 44.2% | 41.3 | 2025-04-10 | |
| 18 | 41.7% | 61.1 | 2025-12-02 | |
| 19 | 41.1% | 53.9 | 2026-02-12 | |
| 20 | 40.1% | 43.4 | 2025-05-20 | |
| 21 | 38.0% | 50.2 | 2025-05-24 | |
| 22 | 30.4% | 40.3 | 2025-05-01 | |
| 23 | 30.4% | – | 2025-11-26 | |
| 24 | 30.1% | 50.4 | 2025-04-17 | |
| 25 | 29.8% | 63.9 | 2026-03-13 | |
| 26 | 28.2% | 65.8 | 2026-09-15 | |
| 27 | 27.6% | 52.9 | 2025-04-10 | |
| 28 | 27.3% | 64.2 | 2026-09-15 | |
| 29 | 26.6% | 67.2 | 2026-09-15 | |
| 30 | 26.4% | – | 2026-09-15 | |
| 31 | 25.7% | 55.0 | 2025-11-26 | |
| 32 | 24.8% | 55.7 | 2025-06-02 | |
| 33 | 24.4% | 49.7 | 2025-04-16 | |
| 34 | 21.6% | – | 2026-09-15 | |
| 35 | 21.5% | 57.4 | 2025-05-06 | |
| 36 | 20.6% | 43.5 | 2025-04-10 | |
| 37 | 20.5% | 55.0 | 2026-02-17 | |
| 38 | 20.2% | 65.7 | 2026-04-08 | |
| 39 | 20.1% | 56.8 | 2026-09-15 | |
| 40 | 19.4% | 53.8 | 2025-04-16 | |
| 41 | 18.2% | 64.2 | 2026-07-08 | |
| 42 | 18.1% | 54.0 | 2025-04-16 | |
| 43 | 17.6% | 44.3 | 2025-08-13 | |
| 44 | 17.5% | 58.4 | 2025-12-15 | |
| 45 | 17.0% | 57.2 | 2025-08-13 | |
| 46 | 17.0% | 56.8 | 2025-08-22 | |
| 47 | 17.0% | 54.3 | 2025-10-02 | |
| 48 | 16.8% | 66.7 | 2026-09-15 | |
| 49 | 16.3% | 67.6 | 2026-07-08 | |
| 50 | 16.1% | 51.9 | 2025-08-08 | |
| 51 | 16.0% | 60.3 | 2025-04-16 | |
| 52 | 15.2% | 60.5 | 2025-11-06 | |
| 53 | 14.8% | 64.3 | 2026-03-23 | |
| 54 | 14.8% | 57.2 | 2025-08-08 | |
| 55 | 13.7% | 64.5 | 2026-09-02 | |
| 56 | 13.6% | 58.2 | 2025-11-26 | |
| 57 | 13.0% | 59.2 | 2026-02-17 | |
| 58 | 13.0% | 46.6 | 2025-06-05 | |
| 59 | 12.8% | 53.4 | 2025-10-02 | |
| 60 | 12.4% | 65.2 | 2026-07-09 |
Cite as: BenchLeader, “FORTRESS leaderboard”, https://www.benchleader.com/benchmarks/scale_fortress, data as of 19 Sept 2026.
FORTRESS: questions
- What does FORTRESS measure?
- Paired adversarial and benign prompts across risk areas; a model must refuse the harmful ones and answer the benign ones. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads FORTRESS?
- DeepSeek R1 leads FORTRESS with 74.4% as of 19 Sept 2026, ahead of GLM-4.5-Air at 63.2%.
- How many models have FORTRESS results?
- 62 model configurations have a FORTRESS result on BenchLeader, all taken from Scale AI SEAL.
- Who runs FORTRESS and how often is it updated?
- FORTRESS is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does FORTRESS count toward the BenchLeader Index?
- No. FORTRESS is shown for reference but left out of the composite index.