Skip to content
Papers.

SAM 3: Segment Anything with Concepts research paper by Meta, 2025

Meta · Nov 20, 2025 · Multimodal and robotics · 999 citations · 138 upvotes · unverified

Read on arXiv

What it shows

Segment Anything Model 3 achieves state-of-the-art performance in promptable concept segmentation and tracking by leveraging a unified model architecture with decoupled recognition and localization.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Meta

All 34
PaperCitations
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable DataWearableQA is a benchmark of multiple-choice questions derived from real longitudinal wearable data that evaluates large language model reasoning across data and health dimensions.Agents and evaluation · Sep 2026 · Unverified0
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and RecipesControlled experiments on when text, image understanding and image generation help or compete when trained together.Multimodal and robotics · Aug 20262
HumanCLAW: Can Vision-Language Models Act Through a Body?HumanCLAW decouples high-level vision-language decisions from low-level motor execution to evaluate embodied action intelligence, revealing that current vision-language models lack embodied self-awareness.Multimodal and robotics · Jul 2026 · Unverified2
TUA-Bench: A Benchmark for General-Purpose Terminal-Use AgentsTUA-Bench presents a comprehensive benchmark for evaluating general-purpose terminal-use agents across diverse digital activities and specialized workflows, revealing significant performance gaps among current frontier agents.Agents and evaluation · Jun 2026 · Unverified2
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved ReasoningProcess-driven image generation decomposes synthesis into iterative steps involving textual planning, visual drafting, textual reflection, and visual refinement, with step-wise supervision ensuring consistency and interpretability.Reasoning · Apr 2026 · Unverified5
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised LearningV-JEPA 2.1 is a self-supervised model that learns dense visual representations for images and videos through a combination of dense predictive loss, deep self-supervision, multi-modal tokenizers, and effective scaling.Multimodal and robotics · Mar 2026 · Unverified83
ReMix: Reinforcement routing for mixtures of LoRAs in LLM finetuningResearchers address imbalance in routing weights of Mixture-of-LoRAs models by proposing Reinforcement Routing (ReMix), which uses non-learnable weights and reinforcement learning techniques to improve model expressiveness and performance.Inference and efficiency · Mar 2026 · Unverified3
Beyond Language Modeling: An Exploration of Multimodal PretrainingControlled multimodal pretraining experiments reveal key insights about unified visual representations, data complementarity, world modeling emergence, and efficient scaling through mixture-of-experts architectures.Multimodal and robotics · Mar 2026 · Unverified32
Topic
PaperCitations
GPT-4o System CardGPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.OpenAI · Oct 2024 · Unverified4,980
DeepSeek-VL: Towards Real-World Vision-Language UnderstandingDeepSeek-VL is an open-source vision-language model that achieves state-of-the-art performance in real-world applications by combining a hybrid vision encoder with effective pretraining strategies to preserve language model capabilities.DeepSeek · Mar 2024 · Unverified889
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and GenerationJanus, an autoregressive framework with separate visual encoding pathways within a unified transformer architecture, enhances performance in unified multimodal understanding and generation.DeepSeek · Oct 2024 · Unverified481
LongLive: Real-time Interactive Long Video GenerationLongLive is a frame-level autoregressive framework for real-time and interactive long video generation, addressing efficiency and quality challenges through causal attention, KV-recache, streaming long tuning, and short window attention.NVIDIA · Sep 2025 · Unverified227
Video models are zero-shot learners and reasonersVeo 3, a generative video model, exhibits zero-shot capabilities across various visual tasks, suggesting a trajectory towards becoming a unified, generalist vision foundation model.Google DeepMind · Sep 2025 · Unverified215
Pixtral 12BPixtral-12B, a 12-billion-parameter multimodal language model, excels in both natural language and image understanding, surpassing larger models and introducing an open-source benchmark for evaluation.Mistral AI · Oct 2024 · Unverified173
About this paper
Authors
Nicolas Carion, Laura Gustafson, Yuan-Ting Hu and 35 more
arXiv
2511.16719 · PDF
Venue
arXiv.org
Citations
999, 153 influential · Semantic Scholar
Upvotes
138 · Hugging Face
Code
github.com/facebookresearch/sam3
Lab
Meta · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 153Sep 25, 2026
Citationsfirst count: 999Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.