Skip to content
Papers.

Direct Preference Optimization: Your Language Model is Secretly a Reward Model research paper by Stanford University, 2023

Stanford University · May 29, 2023 · Alignment and safety · 10,608 citations · 71 upvotes

Read on arXiv

What it shows

DPO aligns a model to human preferences with a simple classification loss, no reward model or RL loop.

Summarised by hand from the abstract.

More from Stanford University

All 2
Topic
About this paper
Authors
Rafael Rafailov, Archit Sharma, Eric Mitchell and 3 more
arXiv
2305.18290 · PDF
Venue
Neural Information Processing Systems
Citations
10,608, 2,203 influential · Semantic Scholar
Upvotes
71 · Hugging Face
Lab
Stanford University

Changes

What changed
New paperAdded to the listSep 24, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.