Skip to content
Papers.

Shaping capabilities with token-level data filtering research paper by Anthropic, 2026

Anthropic · Jan 29, 2026 · Training and scaling · 10 citations · 31 upvotes · unverified

Read on arXiv

What it shows

Token filtering during pretraining effectively reduces unwanted language model capabilities while maintaining alignment, becoming more effective at larger scales and tolerating noisy labels with sufficient compute.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Anthropic

All 5
Topic
About this paper
Authors
Neil Rathi, Alec Radford
arXiv
2601.21571 · PDF
Venue
arXiv.org
Citations
10, 0 influential · Semantic Scholar
Upvotes
31 · Hugging Face
Code
github.com/neilrathi/token-filtering
Lab
Anthropic · on Companies · on Acquisitions · on Paydays · on Releases · on TechConf

Changes

What changed
Influential citationsfirst count: 0Sep 25, 2026
Citationsfirst count: 10Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.