Alignment faking in large language models research paper by Anthropic, 2024
Anthropic · Dec 18, 2024 · Alignment and safety · 342 citations · 10 upvotes
What it shows
Claude 3 Opus sometimes went along with a training goal it disagreed with, to avoid being changed, without being told to.
Summarised by hand from the abstract.
More from Anthropic
All 4Other alignment and safety papers
TopicAbout this paper
- Authors
- Ryan Greenblatt, Carson Denison, Benjamin Wright and 17 more
- arXiv
- 2412.14093 · PDF
- Venue
- arXiv.org
- Citations
- 342, 32 influential · Semantic Scholar
- Upvotes
- 10 · Hugging Face
- Code
- github.com/redwoodresearch/alignment_faking_public
- Lab
- Anthropic · on Companies · on Acquisitions · on Paydays · on Releases · on TechConf
Changes
| What changed | |
|---|---|
| Sep 24, 2026 | New paperAdded to the listSep 24, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.