BenchLeader

FORTRESS

Robustness to harmful requests without over-refusing benign ones. Scale AI.

As of 19 Sept 2026, DeepSeek R1 leads FORTRESS on BenchLeader with 74.4%, ahead of GLM-4.5-Air at 63.2%, across 62 model configurations with a published result.

Published by
Scale AI SEAL
Category
Safety & honesty
Index weight
Reference only
Models
62
Data as of
19 Sept 2026

Scale AI SEAL Leaderboards.

What the test looks like

Paired adversarial and benign prompts across risk areas; a model must refuse the harmful ones and answer the benign ones.

How it is scored

Combined score for harmful-request refusal and benign-request compliance, published by Scale AI.

What to keep in mind

A safety measure, not a capability one; it is not part of the BenchLeader Index.

62 of 62
#
1DeepSeek R1DeepSeekopen74.4%48.72025-04-10
2GLM-4.5-AirZhipu AIopen63.2%47.22025-08-13
3Gemini 2.5 ProGoogle61.7%54.52025-04-14
4Qwen3 235B A22BAlibabaopen61.4%49.62025-07-24
5GLM-4.5Zhipu AIopen59.6%52.12025-08-13
6GPT-4.1 miniOpenAI59.2%45.62025-04-10
7Qwen2.5 72BAlibabaopen56.4%42.72025-04-10
8Mixtral 8x22B InstructMistral AIopen ↗56.1%37.72025-05-23
9Kimi K2Moonshot AIopen55.5%51.02025-07-23
10Gemini 1.5 ProGoogle53.9%42.92025-04-10
11GPT-4.1OpenAI53.0%49.02025-05-23
12Gemini 3.1 Flash LiteGoogle51.0%54.52026-03-13
13Gemini 1.5 FlashGoogle50.6%2025-04-10
14GPT-4o miniOpenAI48.1%36.22025-04-10
15GPT-4oOpenAI47.2%44.32025-04-10
16Llama 3.3 70BMetaopen44.8%41.42025-04-10
17Llama 3.1 70BMetaopen44.2%41.32025-04-10
18Gemini 3 ProGoogle41.7%61.12025-12-02
19Kimi K2.5Moonshot AIopen41.1%53.92026-02-12
20Llama 4 MaverickMetaopen40.1%43.42025-05-20
21Claude 3.7 SonnetAnthropic38.0%50.22025-05-24
22Claude 3.5 HaikuAnthropic30.4%40.32025-05-01
23GPT 5.1 InstantOpenAI30.4%2025-11-26
24o3-miniOpenAI30.1%50.42025-04-17
25Gemini 3.1 ProGoogle29.8%63.92026-03-13
26GLM-5.3maxZhipu AIopen ↗28.2%65.82026-09-15
27Claude Opus 4Anthropic27.6%52.92025-04-10
28Kimi K3Moonshot AIopen ↗27.3%64.22026-09-15
29Kimi K3maxMoonshot AIopen ↗26.6%67.22026-09-15
30Nemotron 3 Ultra Nvfp4highNVIDIA26.4%2026-09-15
31GPT-5.1thinkingOpenAI25.7%55.02025-11-26
32Claude Opus 4thinkingAnthropic24.8%55.72025-06-02
33Claude Sonnet 4Anthropic24.4%49.72025-04-16
34Qwen3.8 2.4T A95BxhighAlibabaopen ↗21.6%2026-09-15
35o4-miniOpenAI21.5%57.42025-05-06
36Llama 3.1 405BMetaopen20.6%43.52025-04-10
37Claude Opus 4.6no reasoningAnthropic20.5%55.02026-02-17
38Muse SparkMeta20.2%65.72026-04-08
39GPT-5.6 SolOpenAI20.1%56.82026-09-15
40o1OpenAI19.4%53.82025-04-16
41Claude Opus 4.8maxAnthropic18.2%64.22026-07-08
42Claude Sonnet 4thinkingAnthropic18.1%54.02025-04-16
43gpt-oss-20bOpenAIopen17.6%44.32025-08-13
44GPT-5.2OpenAI17.5%58.42025-12-15
45GPT-5OpenAI17.0%57.22025-08-13
46GPT-5 miniOpenAI17.0%56.82025-08-22
47Claude Sonnet 4.5Anthropic17.0%54.32025-10-02
48Claude Opus 5Anthropic16.8%66.72026-09-15
49GPT-5.5xhighOpenAI16.3%67.62026-07-08
50Claude Opus 4.1Anthropic16.1%51.92025-08-08
51o3OpenAI16.0%60.32025-04-16
52GPT-5 ProOpenAI15.2%60.52025-11-06
53GPT-5.4 ProOpenAI14.8%64.32026-03-23
54Claude Opus 4.1thinkingAnthropic14.8%57.22025-08-08
55Claude Fable 5.1Anthropic13.7%64.52026-09-02
56Claude Opus 4.5Anthropic13.6%58.22025-11-26
57Claude Opus 4.6maxAnthropic13.0%59.22026-02-17
58Claude 3.5 SonnetAnthropic13.0%46.62025-06-05
59Claude Sonnet 4.5thinkingAnthropic12.8%53.42025-10-02
60Muse Spark 1.1Meta12.4%65.22026-07-09

Cite as: BenchLeader, “FORTRESS leaderboard”, https://www.benchleader.com/benchmarks/scale_fortress, data as of 19 Sept 2026.

FORTRESS: questions

What does FORTRESS measure?
Paired adversarial and benign prompts across risk areas; a model must refuse the harmful ones and answer the benign ones. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads FORTRESS?
DeepSeek R1 leads FORTRESS with 74.4% as of 19 Sept 2026, ahead of GLM-4.5-Air at 63.2%.
How many models have FORTRESS results?
62 model configurations have a FORTRESS result on BenchLeader, all taken from Scale AI SEAL.
Who runs FORTRESS and how often is it updated?
FORTRESS is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
Does FORTRESS count toward the BenchLeader Index?
No. FORTRESS is shown for reference but left out of the composite index.