PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation research paper by NVIDIA, 2026
NVIDIA · Sep 29, 2026 · Multimodal and robotics · 28 upvotes · unverified 6 days ago
What it shows
Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language...
By Cong Wei, Xuanchi Ren, Bryan Chu and 8 more · arXiv 2609.38597 · PDF · Code
UnverifiedHugging Face's summary; not yet checked by hand.
More from NVIDIA
All 79Other multimodal and robotics papers
TopicAbout this paper
- Authors
- Cong Wei, Xuanchi Ren, Bryan Chu and 8 more
- arXiv
- 2609.38597 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 28 · Hugging Face
- Code
- github.com/nv-tlabs/PixelUMM
- Lab
- NVIDIA · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Oct 5, 2026today | New paperFound by the weekly scan, unverifiedOct 5, 2026today |