Skip to content
Papers.

Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming research paper by Anthropic, 2025

Anthropic · Jan 31, 2025 · Alignment and safety · 199 citations · 10 upvotes

Read on arXiv

What it shows

Classifiers trained from a written constitution held off universal jailbreaks through more than 3,000 hours of red teaming.

Summarised by hand from the abstract.

More from Anthropic

All 4
Topic
About this paper
Authors
Mrinank Sharma, Meg Tong, Jesse Mu and 40 more
arXiv
2501.18837 · PDF
Venue
arXiv.org
Citations
199, 20 influential · Semantic Scholar
Upvotes
10 · Hugging Face
Lab
Anthropic · on Companies · on Acquisitions · on Paydays · on Releases · on TechConf

Changes

What changed
New paperAdded to the listSep 24, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.