Skip to content
Papers.

UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling research paper by XPENG Robotics, 2026

XPENG Robotics · Apr 21, 2026 · Multimodal and robotics · 11 citations · 32 upvotes · unverified

Read on arXiv

What it shows

UniT enables human-to-humanoid transfer by creating a unified visual-language representation that bridges kinematic differences through cross-reconstruction mechanisms and shared latent spaces.

UnverifiedHugging Face's summary; not yet checked by hand.

Topic
PaperCitations
Segment AnythingA promptable model and a dataset of over a billion masks that cut out any object in any image.Meta · Apr 202315.7k
GPT-4o System CardGPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.OpenAI · Oct 2024 · Unverified4,980
SAM 3: Segment Anything with ConceptsSegment Anything Model 3 achieves state-of-the-art performance in promptable concept segmentation and tracking by leveraging a unified model architecture with decoupled recognition and localization.Meta · Nov 2025 · Unverified999
DeepSeek-VL: Towards Real-World Vision-Language UnderstandingDeepSeek-VL is an open-source vision-language model that achieves state-of-the-art performance in real-world applications by combining a hybrid vision encoder with effective pretraining strategies to preserve language model capabilities.DeepSeek · Mar 2024 · Unverified889
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and GenerationJanus, an autoregressive framework with separate visual encoding pathways within a unified transformer architecture, enhances performance in unified multimodal understanding and generation.DeepSeek · Oct 2024 · Unverified481
SAM 3D: 3Dfy Anything in ImagesSAM 3D is a generative model that reconstructs 3D objects from single images using a multi-stage training framework that includes synthetic pretraining and real-world alignment, achieving high performance in human preference tests.Meta · Nov 2025 · Unverified256
About this paper
Authors
Boyu Chen, Yi Chen, Lu Qiu and 3 more
arXiv
2604.19734 · PDF
Venue
arXiv.org
Citations
11, 0 influential · Semantic Scholar
Upvotes
32 · Hugging Face
Code
github.com/xpeng-robotics/UniT
Lab
XPENG Robotics · on Companies · on Rounds

Changes

What changed
Influential citationsfirst count: 0Sep 25, 2026
Citationsfirst count: 11Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.