Skip to content
Papers.

Alignment and safety

Making models follow intent and resist misuse: RLHF, preferences, jailbreaks, deception, hallucination. 7 papers from 3 labs.

7 of 7 papers, newest first

PaperCitations
Why Language Models HallucinateModels hallucinate because training and benchmarks reward confident guessing over saying they do not know.OpenAI · Sep 2025329
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red TeamingClassifiers trained from a written constitution held off universal jailbreaks through more than 3,000 hours of red teaming.Anthropic · Jan 2025199
Alignment faking in large language modelsClaude 3 Opus sometimes went along with a training goal it disagreed with, to avoid being changed, without being told to.Anthropic · Dec 2024342
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingModels trained with a hidden backdoor kept their deceptive behaviour through standard safety training.Anthropic · Jan 2024576
Direct Preference Optimization: Your Language Model is Secretly a Reward ModelDPO aligns a model to human preferences with a simple classification loss, no reward model or RL loop.Stanford University · May 202310.6k
Constitutional AI: Harmlessness from AI FeedbackConstitutional AI trains a harmless assistant from AI feedback guided by a short list of written principles, not human harm labels.Anthropic · Dec 20223,718
Training language models to follow instructions with human feedbackInstructGPT: fine-tuning on human feedback made a 1.3B model preferred over the 175B GPT-3.OpenAI · Mar 202224.2k

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.