Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training research paper by Scale AI, 2025
Scale AI · Sep 25, 2025 · Training and scaling · 20 upvotes 1 year ago
What it shows
Rubric-based rewards mitigate reward over-optimization in reinforcement fine-tuning by leveraging off-policy examples while maintaining reward reliability.
Hugging Face's summary; not yet checked by hand.
More from Scale AI
All 25Other training and scaling papers
TopicAbout this paper
- Authors
- Junkai Zhang, Zihao Wang, Lin Gui and 7 more
- arXiv
- 2509.21500 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 20 · Hugging Face
- Code
- github.com/Jun-Kai-Zhang/rubrics
- Lab
- Scale AI · on Companies · on Acquisitions
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |