Training Compute-Optimal Large Language Models research paper by Google DeepMind, 2022
Google DeepMind · Mar 29, 2022 · Training and scaling · 3,755 citations · 13 upvotes
What it shows
Chinchilla: for a fixed compute budget, train a smaller model on more data; parameters and tokens should grow together.
Summarised by hand from the abstract.
More from Google DeepMind
All 3Other training and scaling papers
TopicAbout this paper
- Authors
- Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch and 19 more
- arXiv
- 2203.15556 · PDF
- Venue
- Advances in Neural Information Processing Systems 35
- Citations
- 3,755, 352 influential · Semantic Scholar
- Upvotes
- 13 · Hugging Face
- Code
- github.com/karpathy/llama2.c
- Lab
- Google DeepMind · on Companies · on Acquisitions
Changes
| What changed | |
|---|---|
| Sep 24, 2026 | New paperAdded to the listSep 24, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.