BenchLeader

MASK

Whether a model states what it believes when pressured to lie. Scale AI.

As of 19 Sept 2026, Claude Opus 4.6 leads MASK on BenchLeader with 96.3%, ahead of Claude Sonnet 4.5 at 96.1%, across 64 model configurations with a published result.

Published by
Scale AI SEAL
Category
Safety & honesty
Index weight
Reference only
Models
64
Data as of
19 Sept 2026

Scale AI SEAL Leaderboards.

What the test looks like

MASK separates honesty from accuracy: the model's belief is elicited, then it is pressured to say otherwise, and honesty is whether it holds.

How it is scored

Honesty rate, published by Scale AI with the Center for AI Safety.

What to keep in mind

A safety and alignment measure; not part of the index.

64 of 64
#
1Claude Opus 4.6no reasoningAnthropic96.3%55.02026-02-17
2Claude Sonnet 4.5thinkingAnthropic96.1%53.42025-10-02
3Claude Sonnet 4thinkingAnthropic95.3%54.02025-05-24
4Claude Opus 4.1thinkingAnthropic94.2%57.22025-08-08
5Claude Opus 4.5thinkingAnthropic92.5%60.82025-11-26
6gpt-oss-120bOpenAIopen92.0%46.92025-08-13
7GPT-5.4 ProOpenAI91.7%64.32026-03-23
8GPT-5.4thinkingOpenAI89.7%2026-03-10
9Claude Sonnet 4Anthropic89.3%49.72025-05-23
10Claude Opus 4thinkingAnthropic87.9%55.72025-05-24
11Claude Opus 4.1Anthropic87.4%51.92025-08-08
12Claude Opus 4.5Anthropic87.1%58.22025-11-26
13GPT-5.2OpenAI86.7%58.42025-12-15
14gpt-oss-20bOpenAIopen86.5%44.32025-08-13
15Claude Sonnet 4.5Anthropic86.4%54.32025-10-02
16GPT-5.1thinkingOpenAI86.3%55.02025-11-26
17GPT-5 ProOpenAI86.0%60.52025-11-06
18Claude Opus 4.6maxAnthropic85.4%59.22026-02-17
19o3highOpenAI84.5%53.02025-04-25
20GPT-5 miniOpenAI82.6%56.82025-08-22
21o3mediumOpenAI82.6%53.72025-04-16
22o3-prohighOpenAI82.5%54.52025-06-26
23Claude 3.7 SonnetthinkingAnthropic82.1%51.72025-03-05
24Claude Opus 4Anthropic80.3%52.92025-05-23
25GPT-5OpenAI79.3%57.22025-08-13
26Claude 3 OpusAnthropic79.0%39.62025-03-05
27o4-minihighOpenAI78.6%50.82025-04-25
28o4-minimediumOpenAI72.9%52.42025-04-16
29Claude 3.5 SonnetAnthropic72.3%46.62025-03-05
30Claude 3.7 SonnetAnthropic72.3%50.22025-03-05
31Kimi K2.5Moonshot AIopen70.5%53.92026-02-12
32GPT 5.1 InstantOpenAI63.3%2025-11-26
33o1-proOpenAI61.6%2025-03-05
34GLM-4.5Zhipu AIopen61.5%52.12025-08-13
35Llama 3.1 405BMetaopen61.4%43.52025-03-05
36GPT-4.1 nanoOpenAI61.4%38.62025-04-14
37GLM-4.5-AirZhipu AIopen60.8%47.22025-08-13
38GPT-4oOpenAI60.1%44.32025-03-05
39o1OpenAI59.3%53.82025-03-05
40DeepSeek R1DeepSeekopen57.3%48.72025-03-05
41GPT-4.5OpenAI56.9%50.22025-03-05
42Mistral MagistralMistral AI56.5%2025-06-26
43Qwen3 235B A22BAlibabaopen56.4%49.62025-05-01
44Gemini 2.5 ProGoogle55.9%54.52025-03-05
45Llama 3.2 90BMetaopen ↗54.1%37.62025-03-05
46DeepSeek R1 0528DeepSeekopen53.0%50.22025-05-29
47Llama 3.3 70BMetaopen51.9%41.42025-03-05
48GPT-4.1OpenAI51.1%49.02025-04-15
49GPT-4.1 miniOpenAI50.0%45.62025-04-14
50Llama 4 MaverickMetaopen49.7%43.42025-03-05
51o3-minilowOpenAI49.7%40.42026-03-25
52Gemini 2.0 FlashthinkingGoogle49.5%45.22025-03-05
53Gemini 2.5 FlashGoogle49.1%52.52026-03-25
54Gemini 2.0 FlashGoogle49.1%43.72025-03-05
55o3-minimediumOpenAI48.9%45.62025-03-05
56Gemini 2.0 ProGoogle48.7%47.22025-04-01
57Gemini 3.1 Flash LiteGoogle48.4%54.52026-03-23
58Mistral Large 2Mistral AIopen47.5%41.02025-04-01
59o3-minihighOpenAI46.8%47.72025-04-01
60Kimi K2Moonshot AIopen46.7%51.02025-07-23

Cite as: BenchLeader, “MASK leaderboard”, https://www.benchleader.com/benchmarks/scale_mask, data as of 19 Sept 2026.

MASK: questions

What does MASK measure?
MASK separates honesty from accuracy: the model's belief is elicited, then it is pressured to say otherwise, and honesty is whether it holds. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads MASK?
Claude Opus 4.6 leads MASK with 96.3% as of 19 Sept 2026, ahead of Claude Sonnet 4.5 at 96.1%.
How many models have MASK results?
64 model configurations have a MASK result on BenchLeader, all taken from Scale AI SEAL.
Who runs MASK and how often is it updated?
MASK is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
Does MASK count toward the BenchLeader Index?
No. MASK is shown for reference but left out of the composite index.