Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs research paper by Cohere, 2024
Cohere · Feb 22, 2024 · Alignment and safety · 18 upvotes 2 years ago
What it shows
Reinforcement Learning from Human Feedback (RLHF) benefits from simpler REINFORCE-style optimization methods over computationally expensive PPO or "RL-free" methods, enhancing performance and reducing cost in large language models.
Hugging Face's summary; not yet checked by hand.
More from Cohere
All 30Other alignment and safety papers
TopicAbout this paper
- Authors
- Arash Ahmadian, Chris Cremer, Matthias Gallé and 4 more
- arXiv
- 2402.14740 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 18 · Hugging Face
- Lab
- Cohere · on Companies · on Acquisitions · on Releases
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |