Skip to content

LEGO-Anything: Coding Agents for 3D Scene Reconstruction research paper by AWS, 2026

AWS · Sep 28, 2026 · Multimodal and robotics · 135 upvotes · unverified 7 days ago

Read on arXiv

What it shows

A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program whose execution yields a scene that can be inspected, edited, and...

By Xirui Li, Peng Shi, Mingwen Dong and 7 more · arXiv 2609.36380 · PDF

UnverifiedHugging Face's summary; not yet checked by hand.

More from AWS

All 3
Topic
PaperCitations
Segment AnythingA promptable model and a dataset of over a billion masks that cut out any object in any image.Meta · Apr 20233 years ago15.9k
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsBLIP-2 efficiently pre-trains vision-language models using off-the-shelf frozen encoders and decoders, achieving state-of-the-art performance with fewer parameters.Salesforce · Jan 20233 years ago9,360
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationBLIP, a Vision-Language Pre-training framework, improves performance across both understanding and generation tasks by bootstrapping captions from noisy web data, achieving state-of-the-art results on image-text retrieval, image captioning, and VQA, and showing strong generalization to video-language tasks.Salesforce · Jan 20224 years ago7,410
GPT-4o System CardGPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.OpenAI · Oct 2024 · Unverified1 year ago5,071
SAM 3: Segment Anything with ConceptsSegment Anything Model 3 achieves state-of-the-art performance in promptable concept segmentation and tracking by leveraging a unified model architecture with decoupled recognition and localization.Meta · Nov 2025 · Unverified10 months ago1,105
DeepSeek-VL: Towards Real-World Vision-Language UnderstandingDeepSeek-VL is an open-source vision-language model that achieves state-of-the-art performance in real-world applications by combining a hybrid vision encoder with effective pretraining strategies to preserve language model capabilities.DeepSeek · Mar 2024 · Unverified2 years ago896
About this paper
Authors
Xirui Li, Peng Shi, Mingwen Dong and 7 more
arXiv
2609.36380 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
135 · Hugging Face
Lab
AWS · on Companies · on Acquisitions · on Releases · on TechConf

Changes

What changed
New paperFound by the weekly scan, unverifiedOct 5, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.