Skip to content
Papers.

Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving research paper by Google, 2026

Google · May 21, 2026 · Multimodal and robotics · 1 citations · 25 upvotes · unverified

Read on arXiv

What it shows

Sensor2Sensor generates high-fidelity multi-modal sensor data from in-the-wild dashcam videos using diffusion models and 4D Gaussian Splatting for autonomous driving system training and validation.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Google

All 33
PaperCitations
RRSI: Regularized Recursive Self-Improvement of Agent HarnessesRRSI keeps self-improving agent harnesses from memorising their training tasks, so gains carry over to new benchmarks.Agents and evaluation · Sep 20260
Verifiable Social Reasoning for LLM AssistantsLLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social...Reasoning · Sep 2026 · Unverified0
Dream-RSI: Recursive Self-Improvement through Evolving WorldsDream-RSI enables scalable recursive self-improvement by using historical discovery replay to evaluate exploration policies offline, reducing costly online evaluations.Retrieval and data · Sep 2026 · Unverified0
Procedural Graphs: Self-Evolving Execution Structures for LLM AgentsA procedural graph framework organizes agent actions into structured relational triplets, providing situational guidance and self-evolving topology to improve long-horizon tool use.Foundation models · Sep 2026 · Unverified0
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill EvolutionWikiSkill co-evolves reusable agent skills with a persistent knowledge base to systematically accumulate experience and improve performance across models.Agents and evaluation · Aug 2026 · Unverified0
EnvHarness: Awakening Static Worlds for Agent LearningEnvHarness and EnvRigger dynamically reshape static environments via programmable plugins to target agent weaknesses and improve reinforcement learning co-evolution.Agents and evaluation · Aug 2026 · Unverified2
Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump ProcessesA coupled Markov jump process with cross-modal attention and remasking enables a training-free single-pass sampler for joint multimodal generation that improves with more denoising steps.Multimodal and robotics · Jul 2026 · Unverified1
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMsReinforcement learning with metacognitive feedback and metacognitive data selection improve large language model calibration by enabling accurate self-assessment of performance and uncertainty.Foundation models · Jun 2026 · Unverified1
Topic
PaperCitations
Segment AnythingA promptable model and a dataset of over a billion masks that cut out any object in any image.Meta · Apr 202315.7k
GPT-4o System CardGPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.OpenAI · Oct 2024 · Unverified4,980
SAM 3: Segment Anything with ConceptsSegment Anything Model 3 achieves state-of-the-art performance in promptable concept segmentation and tracking by leveraging a unified model architecture with decoupled recognition and localization.Meta · Nov 2025 · Unverified999
DeepSeek-VL: Towards Real-World Vision-Language UnderstandingDeepSeek-VL is an open-source vision-language model that achieves state-of-the-art performance in real-world applications by combining a hybrid vision encoder with effective pretraining strategies to preserve language model capabilities.DeepSeek · Mar 2024 · Unverified889
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and GenerationJanus, an autoregressive framework with separate visual encoding pathways within a unified transformer architecture, enhances performance in unified multimodal understanding and generation.DeepSeek · Oct 2024 · Unverified481
SAM 3D: 3Dfy Anything in ImagesSAM 3D is a generative model that reconstructs 3D objects from single images using a multi-stage training framework that includes synthetic pretraining and real-world alignment, achieving high performance in human preference tests.Meta · Nov 2025 · Unverified256
About this paper
Authors
Jiahao Wang, Bo Sun, Yijing Bai and 12 more
arXiv
2605.22809 · PDF
Venue
arXiv.org
Citations
1, 0 influential · Semantic Scholar
Upvotes
25 · Hugging Face
Lab
Google · on Companies · on Acquisitions · on Paydays · on TechConf · on Releases

Changes

What changed
Influential citationsfirst count: 0Sep 25, 2026
Citationsfirst count: 1Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.