Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training research paper by Salesforce, 2026
Salesforce · Sep 13, 2026 · Architectures · 19 upvotes 13 days ago
What it shows
Bounds the four memory peaks that crash long-context mixture-of-experts training, from expert dispatch to optimizer state, with fixed-size GPU schedules.
Summarised by hand from the abstract.
More from Salesforce
All 33Other architectures papers
TopicAbout this paper
- Authors
- Shrey Pandit, Xuan-Phi Nguyen, Yiran Zhao and 1 more
- arXiv
- 2609.14306 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 19 · Hugging Face
- Lab
- Salesforce · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |