Reward Hacking in Rubric-Based Reinforcement Learning research paper by Scale AI, 2026
Scale AI · May 12, 2026 · Reasoning · 4 upvotes 4 months ago
What it shows
Research examines reward hacking in rubric-based reinforcement learning, identifying verifier failure and rubric-design limitations as key sources of divergence between training and evaluation metrics.
Hugging Face's summary; not yet checked by hand.
More from Scale AI
All 25Other reasoning papers
TopicAbout this paper
- Authors
- Anas Mahmoud, MohammadHossein Rezaei, Zihao Wang and 3 more
- arXiv
- 2605.12474 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 4 · Hugging Face
- Lab
- Scale AI · on Companies · on Acquisitions
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |