Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains research paper by Scale AI, 2025
Scale AI · Jul 23, 2025 · Agents and evaluation · 6 upvotes 1 year ago
What it shows
Rubrics as Rewards (RaR) framework uses structured rubrics as interpretable reward signals for on-policy training, improving performance in real-world reinforcement learning tasks with subjective criteria.
Hugging Face's summary; not yet checked by hand.
More from Scale AI
All 25Other agents and evaluation papers
TopicAbout this paper
- Authors
- Anisha Gunjal, Anthony Wang, Elaine Lau and 3 more
- arXiv
- 2507.17746 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 6 · Hugging Face
- Lab
- Scale AI · on Companies · on Acquisitions
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |