Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents research paper by Scale AI, 2024
Scale AI · Oct 11, 2024 · Agents and evaluation 1 year ago
What it shows
BrowserART, a test suite for red-teaming browser agents, reveals that LLMs trained to refuse harmful instructions in chats often fail to do so when equipped with web browser capabilities.
Hugging Face's summary; not yet checked by hand.
More from Scale AI
All 25Other agents and evaluation papers
TopicAbout this paper
- Authors
- Priyanshu Kumar, Elaine Lau, Saranya Vijayakumar and 9 more
- arXiv
- 2410.13886 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 0 · Hugging Face
- Code
- github.com/scaleapi/browser-art
- Lab
- Scale AI · on Companies · on Acquisitions
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |