An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning research paper by NVIDIA, 2026
NVIDIA · Sep 28, 2026 · Reasoning · 25 upvotes · unverified 7 days ago
What it shows
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization.
By Shangzhe Li, Yuxiao Yang, Tianrun Yu and 4 more · arXiv 2609.35505 · PDF · Code
UnverifiedHugging Face's summary; not yet checked by hand.
More from NVIDIA
All 79Other reasoning papers
TopicAbout this paper
- Authors
- Shangzhe Li, Yuxiao Yang, Tianrun Yu and 4 more
- arXiv
- 2609.35505 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 25 · Hugging Face
- Code
- github.com/UNCSciML/LSPD
- Lab
- NVIDIA · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Oct 5, 2026today | New paperFound by the weekly scan, unverifiedOct 5, 2026today |