Autoregressive Universal Video Segmentation Model research paper by NVIDIA, 2025
NVIDIA · Aug 26, 2025 · Multimodal and robotics · 1 citations · 29 upvotes · unverified
What it shows
AUSM, an autoregressive universal segmentation model, unifies prompted and unprompted video segmentation by treating it as sequential mask prediction, achieving superior performance and faster training on standard benchmarks.
UnverifiedHugging Face's summary; not yet checked by hand.
More from NVIDIA
All 75Other multimodal and robotics papers
TopicAbout this paper
- Authors
- Miran Heo, Sukjun Hwang, Min-Hung Chen and 4 more
- arXiv
- 2508.19242 · PDF
- Venue
- arXiv.org
- Citations
- 1, 0 influential · Semantic Scholar
- Upvotes
- 29 · Hugging Face
- Lab
- NVIDIA · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Sep 25, 2026 | Upvotes28 to 29 (+1)Sep 25, 2026 |
| Sep 25, 2026 | Influential citationsfirst count: 0Sep 25, 2026 |
| Sep 25, 2026 | Citationsfirst count: 1Sep 25, 2026 |
| Sep 25, 2026 | New paperFound by the weekly scan, unverifiedSep 25, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.