Skip to content
Papers.

LoST: Level of Semantics Tokenization for 3D Shapes research paper by Adobe, 2026

Adobe · Mar 18, 2026 · Multimodal and robotics · 3 citations · 32 upvotes · unverified

Read on arXiv

What it shows

Level-of-Semantics Tokenization (LoST) improves 3D shape generation by ordering tokens based on semantic salience and using a novel relational alignment loss for better reconstruction and efficiency.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Adobe

All 8
PaperCitations
Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic DesignProfessional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle.Agents and evaluation · Sep 2026 · Unverified0
Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image GenerationAtelier improves artist-grounded image generation by translating vague artistic intent into explicit control states that separate scene content from style, reducing reliance on stereotypical shortcuts.Multimodal and robotics · Aug 2026 · Unverified1
WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric RepresentationVideo diffusion transformers enhanced with camera pose representation enable precise action control and long-term 3D consistency in interactive gaming environments through physics-based action spaces and geometric grounding.Multimodal and robotics · Mar 2026 · Unverified10
Memory-V2V: Augmenting Video-to-Video Diffusion Models with MemoryMemory-V2V enhances multi-turn video editing by maintaining cross-consistency through explicit memory mechanisms and efficient token compression in video-to-video diffusion models.Multimodal and robotics · Jan 2026 · Unverified1
Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and EditingLatent diffusion models using representation encoder features face challenges in semantic compactness and pixel-level reconstruction, which are addressed through a semantic-pixel reconstruction objective that enables compact yet semantically rich representations for unified text-to-image and image editing tasks.Multimodal and robotics · Dec 2025 · Unverified22
V-RGBX: Video Editing with Accurate Controls over Intrinsic PropertiesV-RGBX presents an end-to-end framework for intrinsic-aware video editing that combines video inverse rendering, photorealistic video synthesis, and keyframe-based editing with physically grounded intrinsic channel manipulation.Multimodal and robotics · Dec 2025 · Unverified5
MotionStream: Real-Time Video Generation with Interactive Motion ControlsMotionStream enables real-time video generation with sub-second latency and up to 29 FPS by distilling a text-to-video model with motion control into a causal student using Self Forcing with Distribution Matching Distillation and sliding-window causal attention with attention sinks.Multimodal and robotics · Nov 2025 · Unverified76
Topic
PaperCitations
Segment AnythingA promptable model and a dataset of over a billion masks that cut out any object in any image.Meta · Apr 202315.7k
GPT-4o System CardGPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.OpenAI · Oct 2024 · Unverified4,980
SAM 3: Segment Anything with ConceptsSegment Anything Model 3 achieves state-of-the-art performance in promptable concept segmentation and tracking by leveraging a unified model architecture with decoupled recognition and localization.Meta · Nov 2025 · Unverified999
DeepSeek-VL: Towards Real-World Vision-Language UnderstandingDeepSeek-VL is an open-source vision-language model that achieves state-of-the-art performance in real-world applications by combining a hybrid vision encoder with effective pretraining strategies to preserve language model capabilities.DeepSeek · Mar 2024 · Unverified889
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and GenerationJanus, an autoregressive framework with separate visual encoding pathways within a unified transformer architecture, enhances performance in unified multimodal understanding and generation.DeepSeek · Oct 2024 · Unverified481
SAM 3D: 3Dfy Anything in ImagesSAM 3D is a generative model that reconstructs 3D objects from single images using a multi-stage training framework that includes synthetic pretraining and real-world alignment, achieving high performance in human preference tests.Meta · Nov 2025 · Unverified256
About this paper
Authors
Niladri Shekhar Dutt, Zifan Shi, Paul Guerrero and 4 more
arXiv
2603.17995 · PDF
Venue
arXiv.org
Citations
3, 2 influential · Semantic Scholar
Upvotes
32 · Hugging Face
Code
github.com/niladridutt/LoST
Lab
Adobe · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 2Sep 25, 2026
Citationsfirst count: 3Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.