Constitutional AI: Harmlessness from AI Feedback research paper by Anthropic, 2022
Anthropic · Dec 15, 2022 · Alignment and safety · 3,718 citations · 4 upvotes
What it shows
Constitutional AI trains a harmless assistant from AI feedback guided by a short list of written principles, not human harm labels.
Summarised by hand from the abstract.
More from Anthropic
All 4Other alignment and safety papers
TopicAbout this paper
- Authors
- Yuntao Bai, Saurav Kadavath, Sandipan Kundu and 48 more
- arXiv
- 2212.08073 · PDF
- Venue
- arXiv.org
- Citations
- 3,718, 260 influential · Semantic Scholar
- Upvotes
- 4 · Hugging Face
- Lab
- Anthropic · on Companies · on Acquisitions · on Paydays · on Releases · on TechConf
Changes
| What changed | |
|---|---|
| Sep 24, 2026 | New paperAdded to the listSep 24, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.