MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration research paper by Microsoft, 2026
Microsoft · Feb 2, 2026 · Training and scaling · 1 citations · 34 upvotes · unverified
What it shows
Training instability in large language models is linked to weight matrix stable rank decline and Jacobian alignment, which MSign addresses through matrix sign operations to prevent gradient explosions.
UnverifiedHugging Face's summary; not yet checked by hand.
More from Microsoft
All 45Other training and scaling papers
TopicAbout this paper
- Authors
- Lianhai Ren, Yucheng Ding, Xiao Liu and 3 more
- arXiv
- 2602.01734 · PDF
- Venue
- arXiv.org
- Citations
- 1, 0 influential · Semantic Scholar
- Upvotes
- 34 · Hugging Face
- Lab
- Microsoft · on Companies · on Acquisitions · on Quarterly · on Paydays · on Releases · on TechConf
Changes
| What changed | |
|---|---|
| Sep 25, 2026 | Influential citationsfirst count: 0Sep 25, 2026 |
| Sep 25, 2026 | Citationsfirst count: 1Sep 25, 2026 |
| Sep 25, 2026 | New paperFound by the weekly scan, unverifiedSep 25, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.