Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming research paper by Anthropic, 2025
Anthropic · Jan 31, 2025 · Alignment and safety · 199 citations · 10 upvotes
What it shows
Classifiers trained from a written constitution held off universal jailbreaks through more than 3,000 hours of red teaming.
Summarised by hand from the abstract.
More from Anthropic
All 4Other alignment and safety papers
TopicAbout this paper
- Authors
- Mrinank Sharma, Meg Tong, Jesse Mu and 40 more
- arXiv
- 2501.18837 · PDF
- Venue
- arXiv.org
- Citations
- 199, 20 influential · Semantic Scholar
- Upvotes
- 10 · Hugging Face
- Lab
- Anthropic · on Companies · on Acquisitions · on Paydays · on Releases · on TechConf
Changes
| What changed | |
|---|---|
| Sep 24, 2026 | New paperAdded to the listSep 24, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.