LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding research paper by NVIDIA, 2026
NVIDIA · May 26, 2026 · Multimodal and robotics · 16 citations · 142 upvotes · unverified
What it shows
Parallel Box Decoding enables efficient and accurate unified visual grounding and detection by decoding geometric elements as atomic units, improving both throughput and localization quality.
UnverifiedHugging Face's summary; not yet checked by hand.
More from NVIDIA
All 75Other multimodal and robotics papers
TopicAbout this paper
- Authors
- Shihao Wang, Shilong Liu, Yuanguo Kuang and 10 more
- arXiv
- 2605.27365 · PDF
- Venue
- arXiv.org
- Citations
- 16, 1 influential · Semantic Scholar
- Upvotes
- 142 · Hugging Face
- Lab
- NVIDIA · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Sep 25, 2026 | Influential citationsfirst count: 1Sep 25, 2026 |
| Sep 25, 2026 | Citationsfirst count: 16Sep 25, 2026 |
| Sep 25, 2026 | New paperFound by the weekly scan, unverifiedSep 25, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.