Learning from Language Feedback via Variational Policy Distillation research paper by Salesforce, 2026
Salesforce · May 18, 2026 · Reasoning · 12 upvotes 4 months ago
What it shows
Variational Policy Distillation enables reinforcement learning from language feedback by co-evolving teacher and student policies through variational expectation-maximization, overcoming limitations of passive distillation in complex reasoning tasks.
Hugging Face's summary; not yet checked by hand.
More from Salesforce
All 33Other reasoning papers
TopicAbout this paper
- Authors
- Yang Li, Erik Nijkamp, Semih Yavuz and 1 more
- arXiv
- 2605.15113 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 12 · Hugging Face
- Lab
- Salesforce · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |