Skip to content

Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence research paper by Amazon, 2026

Amazon · Sep 28, 2026 · Reasoning · 210 upvotes · unverified 7 days ago

Read on arXiv

What it shows

Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words.

By Xi Xiao, Tianchen Zhao, Youngeun Kim and 10 more · arXiv 2609.34563 · PDF · Code

UnverifiedHugging Face's summary; not yet checked by hand.

More from Amazon

All 8
PaperCitations
Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle TranslationLong-form subtitle translation requires reasoning over discourse and cultural context spanning episodes or entire series, while maintaining consistent terminology and style.Agents and evaluation · Sep 2026 · Unverified6 days ago-
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RLAgent trajectories record what an agent does and what happens next.Agents and evaluation · Sep 2026 · Unverified2 weeks ago0
VIDEOP2R: Video Understanding from Perception to ReasoningVideoP2R, a process-aware reinforcement fine-tuning framework, improves video reasoning and understanding by modeling perception and reasoning separately, achieving state-of-the-art results on multiple benchmarks.Reasoning · Nov 2025 · Unverified10 months ago10
Adaptive Multi-Agent Response Refinement in Conversational SystemsA multi-agent framework enhances conversational quality by refining responses through agents responsible for factuality, personalization, and coherence, outperforming existing methods on challenging datasets.Agents and evaluation · Nov 2025 · Unverified10 months ago3
Chronos-2: From Univariate to Universal ForecastingChronos-2, a pretrained model with a group attention mechanism, achieves state-of-the-art performance in zero-shot univariate, multivariate, and covariate-informed forecasting tasks.Inference and efficiency · Oct 2025 · Unverified11 months ago200
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMsThought templates enhance long-context language models by structuring evidence combination and guiding multi-hop inference, leading to consistent performance improvements across various benchmarks.Reasoning · Oct 2025 · Unverified12 months ago4
TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular ReasoningTaTToo, a novel table-grounded Process Reward Model, enhances tabular reasoning by explicitly addressing table-specific operations and integrating tool-based verification, leading to significant performance improvements over existing PRMs.Reasoning · Oct 2025 · Unverified12 months ago13
Topic
PaperCitations
Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsAsking a model to write out its intermediate steps (chain of thought) sharply improves its math and logic answers.Google · Jan 20224 years ago22.5k
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeekMath 7B improves mathematical reasoning through enhanced data pre-training and Group Relative Policy Optimization, achieving high scores on MATH benchmark without external tools.DeepSeek · Feb 2024 · Unverified2 years ago9,498
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningReinforcement learning taught a model long step-by-step reasoning on par with OpenAI o1.DeepSeek · Jan 20251 year ago5,805
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code IntelligenceDeepSeek-Coder-V2, a Mixture-of-Experts language model, excels in code-specific tasks by enhancing coding and mathematical reasoning capabilities while expanding language support and context length.DeepSeek · Jun 2024 · Unverified2 years ago507
MolmoAct: Action Reasoning Models that can Reason in SpaceAction Reasoning Models (ARMs) integrate perception, planning, and control to enable adaptable and explainable robotic behavior, achieving superior performance across various tasks and settings.Ai2 · Aug 2025 · Unverified1 year ago199
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningThinkAct, a dual-system framework, uses reinforced visual latent planning to enable few-shot adaptation, long-horizon planning, and self-correction in embodied AI tasks by bridging high-level reasoning with low-level action execution.NVIDIA · Jul 2025 · Unverified1 year ago177
About this paper
Authors
Xi Xiao, Tianchen Zhao, Youngeun Kim and 10 more
arXiv
2609.34563 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
210 · Hugging Face
Code
github.com/xixiaouab/ReaLVR-code
Lab
Amazon · on Companies · on Quarterly · on Paydays

Changes

What changed
New paperFound by the weekly scan, unverifiedOct 5, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.