PropensityBench
How readily an agent takes harmful actions when pressured in agentic settings. Scale AI.
As of 19 Sept 2026, Gemini 2.5 Pro leads PropensityBench on BenchLeader with 79.0%, ahead of Gemini 2.0 Flash at 77.8%, across 14 model configurations with a published result.
- Published by
- Scale AI SEAL
- Category
- Safety & honesty
- Index weight
- Reference only
- Models
- 14
- Data as of
- 19 Sept 2026
Scale AI SEAL Leaderboards.
What the test looks like
Agents are placed in scenarios where a harmful shortcut is available and pressure rises; the score is how often they take it.
How it is scored
Propensity score, published by Scale AI.
What to keep in mind
A safety measure where lower propensity is better; not part of the index.
- 1Gemini 2.5 Pro79.0%
- 2Gemini 2.0 Flash77.8%
- 3Qwen3 8B75.2%
- 4Gemini 2.5 Flash68.0%
- 5Llama 3.1 8B66.5%
- 6Llama 3.1 70B55.4%
- 7Gemini 3 Pro52.9%
- 8GPT-4o46.1%
- 9GPT-5.234.4%
- 10o3-mini33.2%
- 11Qwen2.5 32B Instruct22.9%
- 12o4-mini15.8%
- 13Claude Sonnet 412.2%
- 14o310.5%
14 of 14
| # | ||||
|---|---|---|---|---|
| 1 | 79.0% | 54.5 | 2025-10-31 | |
| 2 | 77.8% | 43.7 | 2025-10-31 | |
| 3 | 75.2% | – | 2025-10-31 | |
| 4 | 68.0% | 52.5 | 2025-10-31 | |
| 5 | 66.5% | 36.8 | 2025-10-31 | |
| 6 | 55.4% | 41.3 | 2025-10-31 | |
| 7 | 52.9% | 61.1 | 2026-01-15 | |
| 8 | 46.1% | 44.3 | 2025-10-31 | |
| 9 | 34.4% | 58.4 | 2026-01-15 | |
| 10 | 33.2% | 50.4 | 2025-10-31 | |
| 11 | 22.9% | – | 2025-10-31 | |
| 12 | 15.8% | 57.4 | 2025-10-31 | |
| 13 | 12.2% | 49.7 | 2025-10-31 | |
| 14 | 10.5% | 60.3 | 2025-10-31 |
Cite as: BenchLeader, “PropensityBench leaderboard”, https://www.benchleader.com/benchmarks/scale_propensitybench, data as of 19 Sept 2026.
PropensityBench: questions
- What does PropensityBench measure?
- Agents are placed in scenarios where a harmful shortcut is available and pressure rises; the score is how often they take it. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads PropensityBench?
- Gemini 2.5 Pro leads PropensityBench with 79.0% as of 19 Sept 2026, ahead of Gemini 2.0 Flash at 77.8%.
- How many models have PropensityBench results?
- 14 model configurations have a PropensityBench result on BenchLeader, all taken from Scale AI SEAL.
- Who runs PropensityBench and how often is it updated?
- PropensityBench is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does PropensityBench count toward the BenchLeader Index?
- No. PropensityBench is shown for reference but left out of the composite index.