PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach research paper by Scale AI, 2025
Scale AI · Nov 24, 2025 · Agents and evaluation 10 months ago
What it shows
A new benchmark framework, PropensityBench, evaluates models' likelihood to engage in risky behaviors when equipped with simulated dangerous capabilities, highlighting the need for dynamic propensity assessments in AI safety.
Hugging Face's summary; not yet checked by hand.
More from Scale AI
All 25Other agents and evaluation papers
TopicAbout this paper
- Authors
- Udari Madhushani Sehwag, Shayan Shabihi, Alex McAvoy and 4 more
- arXiv
- 2511.20703 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 0 · Hugging Face
- Code
- github.com/scaleapi/propensity-evaluation
- Lab
- Scale AI · on Companies · on Acquisitions
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |