Skip to content

Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States research paper by Alibaba (Qwen), 2026

Alibaba (Qwen) · Oct 1, 2026 · Agents and evaluation · 84 upvotes · unverified 4 days ago

Read on arXiv

What it shows

Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world.

By Yu Luo, Jiamin Jiang, Yimin Zuo and 9 more · arXiv 2610.01415 · PDF · Code

UnverifiedHugging Face's summary; not yet checked by hand.

More from Alibaba (Qwen)

All 63
PaperCitations
QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon AgentsLarge language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout.Training and scaling · Sep 2026 · Unverified8 days ago-
HappyWorld-BenchEvaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification.Agents and evaluation · Sep 2026 · Unverified2 weeks ago0
One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering AgentsRepository-level software engineering (SWE) comprises heterogeneous task categories, whose progress under pooled agentic reinforcement learning can be uneven: gains in some categories coincide with regressions in...Training and scaling · Sep 2026 · Unverified2 weeks ago0
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual DialogueWe define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model.Multimodal and robotics · Sep 2026 · Unverified2 weeks ago0
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use AgentsComputer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line.Agents and evaluation · Sep 2026 · Unverified2 weeks ago1
CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker DistillationCORE distills compositional ranking judgments from a cross-attentive reranker into an embedding model via synthesized multi-level candidates and a Rank-KL objective, improving compositional retrieval without degrading standard performance.Reasoning · Sep 2026 · Unverified4 weeks ago0
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal EnvironmentsTerminal-Universe reconstructs executable workspaces from agent trajectories to synthesize diverse training tasks and improves post-training performance through supervised fine-tuning.Agents and evaluation · Sep 2026 · Unverified4 weeks ago3
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous DrivingQwen-Drive-1.0 is a vision-language foundation model for autonomous driving that unifies 3D perception, visual question answering, and motion planning via shared representations and staged training.Multimodal and robotics · Aug 2026 · Unverified5 weeks ago7
Topic
PaperCitations
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGemini 2.X model family, including Gemini 2.5 Pro and Flash, offers superior coding, reasoning, and multimodal understanding capabilities across a range of computational efficiencies.Google · Jul 2025 · Unverified1 year ago4,244
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.DeepSeek · Dec 2025 · Unverified10 months ago784
Aya Model: An Instruction Finetuned Open-Access Multilingual Language ModelAya, a multilingual generative language model supporting over 50% lower-resourced languages, excels in both generative and discriminative tasks across 99 languages, and provides extensive evaluations and open-source resources.Cohere · Feb 20242 years ago410
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?WorkArena and BrowserGym evaluate large language model-based agents' ability to perform enterprise software tasks, revealing gaps in current agent capabilities and differences between open and closed-source LLMs.ServiceNow · Mar 20242 years ago387
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable DomainsRubrics as Rewards (RaR) framework uses structured rubrics as interpretable reward signals for on-policy training, improving performance in real-world reinforcement learning tasks with subjective criteria.Scale AI · Jul 20251 year ago343
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?SWE-Bench Pro is a challenging benchmark for coding models, featuring complex, enterprise-level problems that require substantial code modifications, with performance evaluations showing significant limitations in current models.Scale AI · Sep 20251 year ago270
About this paper
Authors
Yu Luo, Jiamin Jiang, Yimin Zuo and 9 more
arXiv
2610.01415 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
84 · Hugging Face
Code
github.com/luoyu100/PoS
Lab
Alibaba (Qwen) · on Companies · on Quarterly

Changes

What changed
New paperFound by the weekly scan, unverifiedOct 5, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.