Shaping capabilities with token-level data filtering research paper by Anthropic, 2026
Anthropic · Jan 29, 2026 · Training and scaling · 10 citations · 31 upvotes · unverified
What it shows
Token filtering during pretraining effectively reduces unwanted language model capabilities while maintaining alignment, becoming more effective at larger scales and tolerating noisy labels with sufficient compute.
UnverifiedHugging Face's summary; not yet checked by hand.
More from Anthropic
All 5Other training and scaling papers
TopicAbout this paper
- Authors
- Neil Rathi, Alec Radford
- arXiv
- 2601.21571 · PDF
- Venue
- arXiv.org
- Citations
- 10, 0 influential · Semantic Scholar
- Upvotes
- 31 · Hugging Face
- Code
- github.com/neilrathi/token-filtering
- Lab
- Anthropic · on Companies · on Acquisitions · on Paydays · on Releases · on TechConf
Changes
| What changed | |
|---|---|
| Sep 25, 2026 | Influential citationsfirst count: 0Sep 25, 2026 |
| Sep 25, 2026 | Citationsfirst count: 10Sep 25, 2026 |
| Sep 25, 2026 | New paperFound by the weekly scan, unverifiedSep 25, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.