Skip to content

LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet research paper by Scale AI, 2024

Scale AI · Aug 27, 2024 · Foundation models 2 years ago

Read on arXiv

What it shows

Multi-turn human jailbreaks reveal significant vulnerabilities in large language model defenses, achieving high attack success rates and exposing flaws in machine unlearning mechanisms.

Hugging Face's summary; not yet checked by hand.

More from Scale AI

All 25
PaperCitations
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent FailuresContinual Search improves automated root-cause attribution in long agent execution traces by iteratively prompting diagnosis until unresolved evidence is found.Agents and evaluation · Sep 20262 weeks ago-
SteerDuplex: Steerable Duplex Speech Dialogue ModelsA full-duplex speech model that follows spoken instructions on tone, persona, pace and voice, with a benchmark of 390 prompts to test it.Multimodal and robotics · Sep 20262 weeks ago-
Studying Without a Syllabus: Task-Agnostic Environment PreprocessingAn agent can explore unfamiliar environments without task-specific guidance to build reusable artifacts that reduce later inference costs, though larger study budgets do not always improve results.Agents and evaluation · Sep 20262 weeks ago-
HarnessOpt-Bench: Evaluating LLMs at Harness OptimizationHarnessOpt-Bench measures how well frontier LLMs iteratively improve agent harnesses under constrained evaluation budgets, revealing substantial variation across models and tasks.Agents and evaluation · Aug 20267 weeks ago-
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent FailuresAn interaction-centric taxonomy localizes agent failures to specific component interactions to guide targeted repairs across diverse architectures.Agents and evaluation · Jul 20268 weeks ago-
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding SessionsSWE-Interact presents a testbed that evaluates coding agents in realistic multi-turn, user-driven software engineering scenarios, revealing significant gaps between single-turn performance and interactive task completion.Agents and evaluation · Jun 20262 months ago-
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVRPOW3R is a policy-aware framework for reinforcement learning with rubric-based rewards that adapts criterion weights during training to improve policy optimization while preserving human-defined criteria importance.Training and scaling · May 20264 months ago-
Reward Hacking in Rubric-Based Reinforcement LearningResearch examines reward hacking in rubric-based reinforcement learning, identifying verifier failure and rubric-design limitations as key sources of divergence between training and evaluation metrics.Reasoning · May 20264 months ago-
Topic
About this paper
Authors
Nathaniel Li, Ziwen Han, Ian Steneker and 6 more
arXiv
2408.15221 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
0 · Hugging Face
Lab
Scale AI · on Companies · on Acquisitions

Changes

What changed
New paperAdded to the listSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.