RISE: Recursive Improvement via Self-Extrapolating Policy Distillation research paper by Salesforce, 2026
Salesforce · Sep 4, 2026 · Training and scaling · 17 upvotes 3 weeks ago
What it shows
RISE improves language model post-training by recursively generating dense token-level supervision from the model's own reinforcement learning trajectory via self-extrapolation, avoiding external teachers.
Hugging Face's summary; not yet checked by hand.
More from Salesforce
All 33Other training and scaling papers
TopicAbout this paper
- Authors
- Yang Li, Semih Yavuz, Shafiq Joty
- arXiv
- 2609.05295 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 17 · Hugging Face
- Lab
- Salesforce · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |