LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet research paper by Scale AI, 2024
Scale AI · Aug 27, 2024 · Foundation models 2 years ago
What it shows
Multi-turn human jailbreaks reveal significant vulnerabilities in large language model defenses, achieving high attack success rates and exposing flaws in machine unlearning mechanisms.
Hugging Face's summary; not yet checked by hand.
More from Scale AI
All 25Other foundation models papers
TopicAbout this paper
- Authors
- Nathaniel Li, Ziwen Han, Ian Steneker and 6 more
- arXiv
- 2408.15221 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 0 · Hugging Face
- Lab
- Scale AI · on Companies · on Acquisitions
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |