Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments research paper by Alibaba (Qwen), 2026
Alibaba (Qwen) · May 28, 2026 · Multimodal and robotics · 37 citations · 145 upvotes · unverified
What it shows
A unified vision-language-action model is presented that integrates diverse embodied decision-making tasks through a shared architecture and training approach, demonstrating strong performance across manipulation, navigation, and trajectory prediction with generalization across different robot platforms and environments.
UnverifiedHugging Face's summary; not yet checked by hand.
More from Alibaba (Qwen)
All 61Other multimodal and robotics papers
TopicAbout this paper
- Authors
- Qiuyue Wang, Mingsheng Li, Jian Guan and 37 more
- arXiv
- 2605.30280 · PDF
- Venue
- arXiv.org
- Citations
- 37, 3 influential · Semantic Scholar
- Upvotes
- 145 · Hugging Face
- Lab
- Alibaba (Qwen) · on Companies · on Quarterly
Changes
| What changed | |
|---|---|
| Sep 25, 2026 | Influential citationsfirst count: 3Sep 25, 2026 |
| Sep 25, 2026 | Citationsfirst count: 37Sep 25, 2026 |
| Sep 25, 2026 | New paperFound by the weekly scan, unverifiedSep 25, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.