Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training research paper by Anthropic, 2024
Anthropic · Jan 10, 2024 · Alignment and safety · 576 citations · 30 upvotes
What it shows
Models trained with a hidden backdoor kept their deceptive behaviour through standard safety training.
Summarised by hand from the abstract.
More from Anthropic
All 4Other alignment and safety papers
TopicAbout this paper
- Authors
- Evan Hubinger, Carson Denison, Jesse Mu and 36 more
- arXiv
- 2401.05566 · PDF
- Venue
- arXiv.org
- Citations
- 576, 52 influential · Semantic Scholar
- Upvotes
- 30 · Hugging Face
- Code
- github.com/anthropics/sleeper-agents-paper
- Lab
- Anthropic · on Companies · on Acquisitions · on Paydays · on Releases · on TechConf
Changes
| What changed | |
|---|---|
| Sep 24, 2026 | New paperAdded to the listSep 24, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.