Alignment and safety
Making models follow intent and resist misuse: RLHF, preferences, jailbreaks, deception, hallucination. 7 papers from 3 labs.
7 of 7 papers, newest first
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.