Sharpening Tax in Post-Training research paper by Meta, 2026
Meta · Oct 1, 2026 · Training and scaling · 85 upvotes · unverified 4 days ago
What it shows
An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of...
By Changdae Oh, Qi Zeng, Qi Qi and 7 more · arXiv 2610.01509 · PDF · Code
UnverifiedHugging Face's summary; not yet checked by hand.
More from Meta
All 37Other training and scaling papers
TopicAbout this paper
- Authors
- Changdae Oh, Qi Zeng, Qi Qi and 7 more
- arXiv
- 2610.01509 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 85 · Hugging Face
- Code
- github.com/changdaeoh/sharpening-tax
- Lab
- Meta · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Oct 5, 2026today | New paperFound by the weekly scan, unverifiedOct 5, 2026today |