Cosmos 3: Omnimodal World Models for Physical AI research paper by NVIDIA, 2026
NVIDIA · Jun 1, 2026 · Multimodal and robotics · 90 citations · 142 upvotes
What it shows
One model that reads and generates text, images, video, audio and robot actions for physical AI.
Summarised by hand from the abstract.
More from NVIDIA
All 3Other multimodal and robotics papers
TopicAbout this paper
- Authors
- Aditi, Niket Agarwal, Arslan Ali and 288 more
- arXiv
- 2606.02800 · PDF
- Venue
- arXiv.org
- Citations
- 90, 9 influential · Semantic Scholar
- Upvotes
- 142 · Hugging Face
- Code
- github.com/NVIDIA/cosmos
- Lab
- NVIDIA · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Sep 24, 2026 | New paperAdded to the listSep 24, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.