Skip to content

Future Optical Flow Prediction Improves Robot Control & Video Generation research paper by Salesforce, 2026

Salesforce · Jan 15, 2026 · Multimodal and robotics · 19 upvotes 8 months ago

Read on arXiv

What it shows

A novel language-conditioned optical flow forecasting model combines Vision-Language Model and Diffusion architecture to predict future motion from noisy web-scale video data, demonstrating versatility in robotic manipulation and video generation tasks.

Hugging Face's summary; not yet checked by hand.

More from Salesforce

All 33
PaperCitations
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM AgentsKeeps an agent's raw past runs and writes a task-specific memory only when a new task arrives, rather than deciding up front what to keep.Agents and evaluation · Sep 20263 days ago-
Flattening Every Memory Peak in Long-Context Mixture-of-Experts TrainingBounds the four memory peaks that crash long-context mixture-of-experts training, from expert dispatch to optimizer state, with fixed-size GPU schedules.Architectures · Sep 202613 days ago-
RISE: Recursive Improvement via Self-Extrapolating Policy DistillationRISE improves language model post-training by recursively generating dense token-level supervision from the model's own reinforcement learning trajectory via self-extrapolation, avoiding external teachers.Training and scaling · Sep 20263 weeks ago-
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?The study introduces a benchmark to evaluate LLM agents under evolving tool, skill, and agent harnesses, revealing persistent gaps in retention, adaptation, and harness-induced forgetting.Agents and evaluation · Sep 20263 weeks ago-
Random Attention: Rethinking KV Cache Eviction for Efficient ReasoningRandom eviction of reasoning tokens matches selective KV cache compression because reasoning traces are self-protecting through redundancy, making scoring unnecessary once prompts are preserved.Reasoning · Sep 20263 weeks ago-
DarwinX: Evolving Agent Harnesses Through Natural SelectionDarwinX evolves agent harnesses via population selection with frozen models, improving verified performance across benchmarks without benchmark-specific patches.Agents and evaluation · Jul 20268 weeks ago-
StateAct: Program State, before Pixels, for Long-Horizon Computer-Use AgentsStateAct improves computer-use agents by grounding actions, verification, and memory in direct program state rather than screenshots, boosting success rates while reducing cost.Agents and evaluation · Jul 20262 months ago-
Evidence-Backed Video Question AnsweringEvidence-Backed Video Question Answering requires models to provide answers with precise spatio-temporal segmentation evidence, revealing gaps between reasoning and visual grounding that are improved by large-scale instruction tuning.Multimodal and robotics · Jul 20262 months ago-
Topic
PaperCitations
Segment AnythingA promptable model and a dataset of over a billion masks that cut out any object in any image.Meta · Apr 20233 years ago15.7k
GPT-4o System CardGPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.OpenAI · Oct 2024 · Unverified1 year ago4,980
SAM 3: Segment Anything with ConceptsSegment Anything Model 3 achieves state-of-the-art performance in promptable concept segmentation and tracking by leveraging a unified model architecture with decoupled recognition and localization.Meta · Nov 2025 · Unverified10 months ago999
DeepSeek-VL: Towards Real-World Vision-Language UnderstandingDeepSeek-VL is an open-source vision-language model that achieves state-of-the-art performance in real-world applications by combining a hybrid vision encoder with effective pretraining strategies to preserve language model capabilities.DeepSeek · Mar 2024 · Unverified2 years ago889
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and GenerationJanus, an autoregressive framework with separate visual encoding pathways within a unified transformer architecture, enhances performance in unified multimodal understanding and generation.DeepSeek · Oct 2024 · Unverified1 year ago481
SAM 3D: 3Dfy Anything in ImagesSAM 3D is a generative model that reconstructs 3D objects from single images using a multi-stage training framework that includes synthetic pretraining and real-world alignment, achieving high performance in human preference tests.Meta · Nov 2025 · Unverified10 months ago256
About this paper
Authors
Kanchana Ranasinghe, Honglu Zhou, Yu Fang and 7 more
arXiv
2601.10781 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
19 · Hugging Face
Code
github.com/SalesforceAIResearch/FOFPred
Lab
Salesforce · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
New paperAdded to the listSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.