When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale research paper by Cohere, 2023
Cohere · Sep 8, 2023 · Training and scaling · 17 upvotes 3 years ago
What it shows
Perplexity outperforms more complex methods for data quality estimation in pruning large language model training datasets, resulting in performance improvement with reduced data.
Hugging Face's summary; not yet checked by hand.
More from Cohere
All 30Other training and scaling papers
TopicAbout this paper
- Authors
- Max Marion, Ahmet Üstün, Luiza Pozzobon and 3 more
- arXiv
- 2309.04564 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 17 · Hugging Face
- Lab
- Cohere · on Companies · on Acquisitions · on Releases
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |