Skip to content
Papers.

Multimodal and robotics

Vision, video, audio, world models and robot control. 4 papers from 3 labs.

4 of 4 papers, newest first

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.