Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism research paper by NVIDIA, 2019
NVIDIA · Sep 17, 2019 · Training and scaling · 3,177 citations · 5 upvotes
What it shows
Megatron-LM splits each Transformer layer across GPUs to train models with billions of parameters.
Summarised by hand from the abstract.
More from NVIDIA
All 3Other training and scaling papers
TopicAbout this paper
- Authors
- Mohammad Shoeybi, Mostofa Patwary, Raul Puri and 3 more
- arXiv
- 1909.08053 · PDF
- Venue
- arXiv.org
- Citations
- 3,177, 393 influential · Semantic Scholar
- Upvotes
- 5 · Hugging Face
- Lab
- NVIDIA · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Sep 24, 2026 | New paperAdded to the listSep 24, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.