Self-Improving Robust Preference Optimization research paper by Cohere, 2024
Cohere · Jun 3, 2024 · Alignment and safety · 19 upvotes 2 years ago
What it shows
SRPO, a self-improving offline RLHF framework, achieves robustness to out-of-distribution tasks by optimizing a min-max objective that jointly enhances self-improvement and generative policies, leading to superior performance compared to DPO.
Hugging Face's summary; not yet checked by hand.
More from Cohere
All 30Other alignment and safety papers
TopicAbout this paper
- Authors
- Eugene Choi, Arash Ahmadian, Matthieu Geist and 2 more
- arXiv
- 2406.01660 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 19 · Hugging Face
- Lab
- Cohere · on Companies · on Acquisitions · on Releases
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |