MASK
Whether a model states what it believes when pressured to lie. Scale AI.
As of 19 Sept 2026, Claude Opus 4.6 leads MASK on BenchLeader with 96.3%, ahead of Claude Sonnet 4.5 at 96.1%, across 64 model configurations with a published result.
- Published by
- Scale AI SEAL
- Category
- Safety & honesty
- Index weight
- Reference only
- Models
- 64
- Data as of
- 19 Sept 2026
Scale AI SEAL Leaderboards.
What the test looks like
MASK separates honesty from accuracy: the model's belief is elicited, then it is pressured to say otherwise, and honesty is whether it holds.
How it is scored
Honesty rate, published by Scale AI with the Center for AI Safety.
What to keep in mind
A safety and alignment measure; not part of the index.
- 1Claude Opus 4.6 (none)96.3%
- 2Claude Sonnet 4.5 (thinking)96.1%
- 3Claude Sonnet 4 (thinking)95.3%
- 4Claude Opus 4.1 (thinking)94.2%
- 5Claude Opus 4.5 (thinking)92.5%
- 6gpt-oss-120b92.0%
- 7GPT-5.4 Pro91.7%
- 8GPT-5.4 (thinking)89.7%
- 9Claude Sonnet 489.3%
- 10Claude Opus 4 (thinking)87.9%
- 11Claude Opus 4.187.4%
- 12Claude Opus 4.587.1%
- 13GPT-5.286.7%
- 14gpt-oss-20b86.5%
- 15Claude Sonnet 4.586.4%
64 of 64
| # | ||||
|---|---|---|---|---|
| 1 | 96.3% | 55.0 | 2026-02-17 | |
| 2 | 96.1% | 53.4 | 2025-10-02 | |
| 3 | 95.3% | 54.0 | 2025-05-24 | |
| 4 | 94.2% | 57.2 | 2025-08-08 | |
| 5 | 92.5% | 60.8 | 2025-11-26 | |
| 6 | 92.0% | 46.9 | 2025-08-13 | |
| 7 | 91.7% | 64.3 | 2026-03-23 | |
| 8 | 89.7% | – | 2026-03-10 | |
| 9 | 89.3% | 49.7 | 2025-05-23 | |
| 10 | 87.9% | 55.7 | 2025-05-24 | |
| 11 | 87.4% | 51.9 | 2025-08-08 | |
| 12 | 87.1% | 58.2 | 2025-11-26 | |
| 13 | 86.7% | 58.4 | 2025-12-15 | |
| 14 | 86.5% | 44.3 | 2025-08-13 | |
| 15 | 86.4% | 54.3 | 2025-10-02 | |
| 16 | 86.3% | 55.0 | 2025-11-26 | |
| 17 | 86.0% | 60.5 | 2025-11-06 | |
| 18 | 85.4% | 59.2 | 2026-02-17 | |
| 19 | 84.5% | 53.0 | 2025-04-25 | |
| 20 | 82.6% | 56.8 | 2025-08-22 | |
| 21 | 82.6% | 53.7 | 2025-04-16 | |
| 22 | 82.5% | 54.5 | 2025-06-26 | |
| 23 | 82.1% | 51.7 | 2025-03-05 | |
| 24 | 80.3% | 52.9 | 2025-05-23 | |
| 25 | 79.3% | 57.2 | 2025-08-13 | |
| 26 | 79.0% | 39.6 | 2025-03-05 | |
| 27 | 78.6% | 50.8 | 2025-04-25 | |
| 28 | 72.9% | 52.4 | 2025-04-16 | |
| 29 | 72.3% | 46.6 | 2025-03-05 | |
| 30 | 72.3% | 50.2 | 2025-03-05 | |
| 31 | 70.5% | 53.9 | 2026-02-12 | |
| 32 | 63.3% | – | 2025-11-26 | |
| 33 | 61.6% | – | 2025-03-05 | |
| 34 | 61.5% | 52.1 | 2025-08-13 | |
| 35 | 61.4% | 43.5 | 2025-03-05 | |
| 36 | 61.4% | 38.6 | 2025-04-14 | |
| 37 | 60.8% | 47.2 | 2025-08-13 | |
| 38 | 60.1% | 44.3 | 2025-03-05 | |
| 39 | 59.3% | 53.8 | 2025-03-05 | |
| 40 | 57.3% | 48.7 | 2025-03-05 | |
| 41 | 56.9% | 50.2 | 2025-03-05 | |
| 42 | 56.5% | – | 2025-06-26 | |
| 43 | 56.4% | 49.6 | 2025-05-01 | |
| 44 | 55.9% | 54.5 | 2025-03-05 | |
| 45 | 54.1% | 37.6 | 2025-03-05 | |
| 46 | 53.0% | 50.2 | 2025-05-29 | |
| 47 | 51.9% | 41.4 | 2025-03-05 | |
| 48 | 51.1% | 49.0 | 2025-04-15 | |
| 49 | 50.0% | 45.6 | 2025-04-14 | |
| 50 | 49.7% | 43.4 | 2025-03-05 | |
| 51 | 49.7% | 40.4 | 2026-03-25 | |
| 52 | 49.5% | 45.2 | 2025-03-05 | |
| 53 | 49.1% | 52.5 | 2026-03-25 | |
| 54 | 49.1% | 43.7 | 2025-03-05 | |
| 55 | 48.9% | 45.6 | 2025-03-05 | |
| 56 | 48.7% | 47.2 | 2025-04-01 | |
| 57 | 48.4% | 54.5 | 2026-03-23 | |
| 58 | 47.5% | 41.0 | 2025-04-01 | |
| 59 | 46.8% | 47.7 | 2025-04-01 | |
| 60 | 46.7% | 51.0 | 2025-07-23 |
Cite as: BenchLeader, “MASK leaderboard”, https://www.benchleader.com/benchmarks/scale_mask, data as of 19 Sept 2026.
MASK: questions
- What does MASK measure?
- MASK separates honesty from accuracy: the model's belief is elicited, then it is pressured to say otherwise, and honesty is whether it holds. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads MASK?
- Claude Opus 4.6 leads MASK with 96.3% as of 19 Sept 2026, ahead of Claude Sonnet 4.5 at 96.1%.
- How many models have MASK results?
- 64 model configurations have a MASK result on BenchLeader, all taken from Scale AI SEAL.
- Who runs MASK and how often is it updated?
- MASK is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does MASK count toward the BenchLeader Index?
- No. MASK is shown for reference but left out of the composite index.