Skip to content
Papers.

LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory research paper by Google DeepMind, 2026

Google DeepMind · Mar 3, 2026 · Reasoning · 35 citations · 63 upvotes · unverified

Read on arXiv

What it shows

LoGeR enables long-term 3D video reconstruction by combining bidirectional priors with a hybrid memory system that includes parametric Test-Time Training and non-parametric sliding window attention mechanisms.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Google DeepMind

All 12
PaperCitations
DiffusionGemma Technical ReportDiffusionGemma is a fine-tuned mixture-of-experts language model that uses discrete diffusion to generate text blocks in parallel, achieving high speed while preserving capabilities like multimodal inputs and reasoning.Foundation models · Jul 2026 · Unverified1
Gemma 4 Technical ReportOpen multimodal models from 2.3B to 31B parameters, with a thinking mode and image and audio input.Foundation models · Jul 2026110
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMsLiteFrame, a lightweight video encoder with Compressed Token Distillation training method, reduces latency and increases frame processing capacity for long-form video understanding in Video LLMs while maintaining accuracy.Inference and efficiency · May 2026 · Unverified1
Understanding the Challenges in Iterative Generative Optimization with LLMsGenerative optimization using large language models faces challenges due to implicit design decisions about artifact modification and learning evidence that significantly impact success across different applications.Foundation models · Mar 2026 · Unverified8
SIMA 2: A Generalist Embodied Agent for Virtual WorldsSIMA 2, built on a Gemini foundation model, interacts in 3D virtual worlds, reasons about goals, handles complex instructions, and autonomously learns new skills through open-ended self-improvement.Agents and evaluation · Dec 2025 · Unverified18
Robot Learning from a Physical World ModelPhysWorld integrates video generation and physical world modeling to enable accurate robotic manipulation from visual demonstrations without real robot data.Multimodal and robotics · Nov 2025 · Unverified21
Vibe Checker: Aligning Code Evaluation with Human PreferenceVibe Checker evaluates LLMs by combining functional correctness and instruction following to better align with human coding preferences.Agents and evaluation · Oct 2025 · Unverified2
Video models are zero-shot learners and reasonersVeo 3, a generative video model, exhibits zero-shot capabilities across various visual tasks, suggesting a trajectory towards becoming a unified, generalist vision foundation model.Multimodal and robotics · Sep 2025 · Unverified215
Topic
PaperCitations
Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsAsking a model to write out its intermediate steps (chain of thought) sharply improves its math and logic answers.Google · Jan 202222.1k
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeekMath 7B improves mathematical reasoning through enhanced data pre-training and Group Relative Policy Optimization, achieving high scores on MATH benchmark without external tools.DeepSeek · Feb 2024 · Unverified8,994
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningReinforcement learning taught a model long step-by-step reasoning on par with OpenAI o1.DeepSeek · Jan 20255,694
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code IntelligenceDeepSeek-Coder-V2, a Mixture-of-Experts language model, excels in code-specific tasks by enhancing coding and mathematical reasoning capabilities while expanding language support and context length.DeepSeek · Jun 2024 · Unverified502
MolmoAct: Action Reasoning Models that can Reason in SpaceAction Reasoning Models (ARMs) integrate perception, planning, and control to enable adaptable and explainable robotic behavior, achieving superior performance across various tasks and settings.Ai2 · Aug 2025 · Unverified185
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningThinkAct, a dual-system framework, uses reinforced visual latent planning to enable few-shot adaptation, long-horizon planning, and self-correction in embodied AI tasks by bridging high-level reasoning with low-level action execution.NVIDIA · Jul 2025 · Unverified172
About this paper
Authors
Junyi Zhang, Charles Herrmann, Junhwa Hur and 5 more
arXiv
2603.03269 · PDF
Venue
arXiv.org
Citations
35, 4 influential · Semantic Scholar
Upvotes
63 · Hugging Face
Code
github.com/Junyi42/LoGeR
Lab
Google DeepMind · on Companies · on Acquisitions

Changes

What changed
Influential citationsfirst count: 4Sep 25, 2026
Citationsfirst count: 35Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.