Skip to content
Papers.

SAM 3D: 3Dfy Anything in Images research paper by Meta, 2025

Meta · Nov 20, 2025 · Multimodal and robotics · 256 citations · 117 upvotes · unverified

Read on arXiv

What it shows

SAM 3D is a generative model that reconstructs 3D objects from single images using a multi-stage training framework that includes synthetic pretraining and real-world alignment, achieving high performance in human preference tests.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Meta

All 34
PaperCitations
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable DataWearableQA is a benchmark of multiple-choice questions derived from real longitudinal wearable data that evaluates large language model reasoning across data and health dimensions.Agents and evaluation · Sep 2026 · Unverified0
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and RecipesControlled experiments on when text, image understanding and image generation help or compete when trained together.Multimodal and robotics · Aug 20262
HumanCLAW: Can Vision-Language Models Act Through a Body?HumanCLAW decouples high-level vision-language decisions from low-level motor execution to evaluate embodied action intelligence, revealing that current vision-language models lack embodied self-awareness.Multimodal and robotics · Jul 2026 · Unverified2
TUA-Bench: A Benchmark for General-Purpose Terminal-Use AgentsTUA-Bench presents a comprehensive benchmark for evaluating general-purpose terminal-use agents across diverse digital activities and specialized workflows, revealing significant performance gaps among current frontier agents.Agents and evaluation · Jun 2026 · Unverified2
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved ReasoningProcess-driven image generation decomposes synthesis into iterative steps involving textual planning, visual drafting, textual reflection, and visual refinement, with step-wise supervision ensuring consistency and interpretability.Reasoning · Apr 2026 · Unverified5
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised LearningV-JEPA 2.1 is a self-supervised model that learns dense visual representations for images and videos through a combination of dense predictive loss, deep self-supervision, multi-modal tokenizers, and effective scaling.Multimodal and robotics · Mar 2026 · Unverified83
ReMix: Reinforcement routing for mixtures of LoRAs in LLM finetuningResearchers address imbalance in routing weights of Mixture-of-LoRAs models by proposing Reinforcement Routing (ReMix), which uses non-learnable weights and reinforcement learning techniques to improve model expressiveness and performance.Inference and efficiency · Mar 2026 · Unverified3
Beyond Language Modeling: An Exploration of Multimodal PretrainingControlled multimodal pretraining experiments reveal key insights about unified visual representations, data complementarity, world modeling emergence, and efficient scaling through mixture-of-experts architectures.Multimodal and robotics · Mar 2026 · Unverified32
Topic
PaperCitations
GPT-4o System CardGPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.OpenAI · Oct 2024 · Unverified4,980
DeepSeek-VL: Towards Real-World Vision-Language UnderstandingDeepSeek-VL is an open-source vision-language model that achieves state-of-the-art performance in real-world applications by combining a hybrid vision encoder with effective pretraining strategies to preserve language model capabilities.DeepSeek · Mar 2024 · Unverified889
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and GenerationJanus, an autoregressive framework with separate visual encoding pathways within a unified transformer architecture, enhances performance in unified multimodal understanding and generation.DeepSeek · Oct 2024 · Unverified481
LongLive: Real-time Interactive Long Video GenerationLongLive is a frame-level autoregressive framework for real-time and interactive long video generation, addressing efficiency and quality challenges through causal attention, KV-recache, streaming long tuning, and short window attention.NVIDIA · Sep 2025 · Unverified227
Video models are zero-shot learners and reasonersVeo 3, a generative video model, exhibits zero-shot capabilities across various visual tasks, suggesting a trajectory towards becoming a unified, generalist vision foundation model.Google DeepMind · Sep 2025 · Unverified215
Pixtral 12BPixtral-12B, a 12-billion-parameter multimodal language model, excels in both natural language and image understanding, surpassing larger models and introducing an open-source benchmark for evaluation.Mistral AI · Oct 2024 · Unverified173
About this paper
Authors
SAM 3D Team, Xingyu Chen, Fu-Jen Chu and 20 more
arXiv
2511.16624 · PDF
Venue
arXiv.org
Citations
256, 57 influential · Semantic Scholar
Upvotes
117 · Hugging Face
Code
github.com/facebookresearch/sam-3d-objects
Lab
Meta · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 57Sep 25, 2026
Citationsfirst count: 256Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.