BenchLeader

PropensityBench

How readily an agent takes harmful actions when pressured in agentic settings. Scale AI.

As of 19 Sept 2026, Gemini 2.5 Pro leads PropensityBench on BenchLeader with 79.0%, ahead of Gemini 2.0 Flash at 77.8%, across 14 model configurations with a published result.

Published by
Scale AI SEAL
Category
Safety & honesty
Index weight
Reference only
Models
14
Data as of
19 Sept 2026

Scale AI SEAL Leaderboards.

What the test looks like

Agents are placed in scenarios where a harmful shortcut is available and pressure rises; the score is how often they take it.

How it is scored

Propensity score, published by Scale AI.

What to keep in mind

A safety measure where lower propensity is better; not part of the index.

14 of 14
#
1Gemini 2.5 ProGoogle79.0%54.52025-10-31
2Gemini 2.0 FlashGoogle77.8%43.72025-10-31
3Qwen3 8BAlibaba75.2%2025-10-31
4Gemini 2.5 FlashGoogle68.0%52.52025-10-31
5Llama 3.1 8BMetaopen66.5%36.82025-10-31
6Llama 3.1 70BMetaopen55.4%41.32025-10-31
7Gemini 3 ProGoogle52.9%61.12026-01-15
8GPT-4oOpenAI46.1%44.32025-10-31
9GPT-5.2OpenAI34.4%58.42026-01-15
10o3-miniOpenAI33.2%50.42025-10-31
11Qwen2.5 32B InstructAlibabaopen ↗22.9%2025-10-31
12o4-miniOpenAI15.8%57.42025-10-31
13Claude Sonnet 4Anthropic12.2%49.72025-10-31
14o3OpenAI10.5%60.32025-10-31

Cite as: BenchLeader, “PropensityBench leaderboard”, https://www.benchleader.com/benchmarks/scale_propensitybench, data as of 19 Sept 2026.

PropensityBench: questions

What does PropensityBench measure?
Agents are placed in scenarios where a harmful shortcut is available and pressure rises; the score is how often they take it. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads PropensityBench?
Gemini 2.5 Pro leads PropensityBench with 79.0% as of 19 Sept 2026, ahead of Gemini 2.0 Flash at 77.8%.
How many models have PropensityBench results?
14 model configurations have a PropensityBench result on BenchLeader, all taken from Scale AI SEAL.
Who runs PropensityBench and how often is it updated?
PropensityBench is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
Does PropensityBench count toward the BenchLeader Index?
No. PropensityBench is shown for reference but left out of the composite index.