Skip to content
Papers.

WildDet3D: Scaling Promptable 3D Detection in the Wild research paper by Ai2, 2026

Ai2 · Apr 9, 2026 · Multimodal and robotics · 11 citations · 95 upvotes · unverified

Read on arXiv

What it shows

A unified 3D object detection framework with a large-scale dataset enables open-world detection with multiple prompt types and geometric cue integration.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Ai2

All 7
PaperCitations
MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction3D point motion forecasting model predicts object trajectories from visual history and language goals, demonstrating superior performance on benchmarks and transferring effectively to robot manipulation and video generation tasks.Multimodal and robotics · Jun 2026 · Unverified4
MolmoAct2: Action Reasoning Models for Real-world DeploymentA fully open robot action model, released with new datasets including 720 hours of two-arm teleoperation.Multimodal and robotics · May 202644
TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for RoboticsTOPReward is a probabilistically grounded temporal value function that uses pretrained video Vision-Language Models to estimate robotic task progress through internal token logits, achieving superior performance in zero-shot evaluations across diverse real-world tasks.Inference and efficiency · Feb 2026 · Unverified28
Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative ReasoningRD-VLA introduces a recurrent architecture for vision-language-action models that adapts computational depth through latent iterative refinement, achieving constant memory usage and improved task success rates.Reasoning · Feb 2026 · Unverified16
Olmo 3Olmo 3, a family of state-of-the-art fully-open language models at 7B and 32B parameter scales, excels in long-context reasoning, function calling, coding, instruction following, general chat, and knowledge recall.Reasoning · Dec 2025 · Unverified0
MolmoAct: Action Reasoning Models that can Reason in SpaceAction Reasoning Models (ARMs) integrate perception, planning, and control to enable adaptable and explainable robotic behavior, achieving superior performance across various tasks and settings.Reasoning · Aug 2025 · Unverified185
Topic
PaperCitations
Segment AnythingA promptable model and a dataset of over a billion masks that cut out any object in any image.Meta · Apr 202315.7k
GPT-4o System CardGPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.OpenAI · Oct 2024 · Unverified4,980
SAM 3: Segment Anything with ConceptsSegment Anything Model 3 achieves state-of-the-art performance in promptable concept segmentation and tracking by leveraging a unified model architecture with decoupled recognition and localization.Meta · Nov 2025 · Unverified999
DeepSeek-VL: Towards Real-World Vision-Language UnderstandingDeepSeek-VL is an open-source vision-language model that achieves state-of-the-art performance in real-world applications by combining a hybrid vision encoder with effective pretraining strategies to preserve language model capabilities.DeepSeek · Mar 2024 · Unverified889
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and GenerationJanus, an autoregressive framework with separate visual encoding pathways within a unified transformer architecture, enhances performance in unified multimodal understanding and generation.DeepSeek · Oct 2024 · Unverified481
SAM 3D: 3Dfy Anything in ImagesSAM 3D is a generative model that reconstructs 3D objects from single images using a multi-stage training framework that includes synthetic pretraining and real-world alignment, achieving high performance in human preference tests.Meta · Nov 2025 · Unverified256
About this paper
Authors
Weikai Huang, Jieyu Zhang, Sijun Li and 14 more
arXiv
2604.08626 · PDF
Venue
arXiv.org
Citations
11, 3 influential · Semantic Scholar
Upvotes
95 · Hugging Face
Code
github.com/allenai/WildDet3D
Lab
Ai2 · on Companies

Changes

What changed
Influential citationsfirst count: 3Sep 25, 2026
Citationsfirst count: 11Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.