Selecting Diverse SFT Traces Improves Post-RL Generalization research paper by Google, 2026
Google · Sep 27, 2026 · Reasoning · 38 upvotes · unverified 8 days ago
What it shows
Verified solutions are not equally useful for preparing reasoning models for reinforcement learning (RL).
By Dylan Zhang, Mingyuan Wu, Jinning Li · arXiv 2609.33780 · PDF
UnverifiedHugging Face's summary; not yet checked by hand.
More from Google
All 36Other reasoning papers
TopicAbout this paper
- Authors
- Dylan Zhang, Mingyuan Wu, Jinning Li
- arXiv
- 2609.33780 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 38 · Hugging Face
- Lab
- Google · on Companies · on Acquisitions · on Paydays · on TechConf · on Releases
Changes
| What changed | |
|---|---|
| Oct 5, 2026today | New paperFound by the weekly scan, unverifiedOct 5, 2026today |