Skip to content
Papers.

AgentFold: Long-Horizon Web Agents with Proactive Context Management research paper by Alibaba (Qwen), 2025

Alibaba (Qwen) · Oct 28, 2025 · Agents and evaluation · 77 citations · 73 upvotes · unverified

Read on arXiv

What it shows

AgentFold, a novel proactive context management paradigm, enhances long-horizon task performance through dynamic context folding, achieving superior results on benchmarks compared to larger models and proprietary agents.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Alibaba (Qwen)

All 61
PaperCitations
HappyWorld-BenchEvaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification.Agents and evaluation · Sep 2026 · Unverified0
One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering AgentsRepository-level software engineering (SWE) comprises heterogeneous task categories, whose progress under pooled agentic reinforcement learning can be uneven: gains in some categories coincide with regressions in...Training and scaling · Sep 2026 · Unverified0
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual DialogueWe define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model.Multimodal and robotics · Sep 2026 · Unverified0
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use AgentsComputer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line.Agents and evaluation · Sep 2026 · Unverified0
CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker DistillationCORE distills compositional ranking judgments from a cross-attentive reranker into an embedding model via synthesized multi-level candidates and a Rank-KL objective, improving compositional retrieval without degrading standard performance.Reasoning · Sep 2026 · Unverified0
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal EnvironmentsTerminal-Universe reconstructs executable workspaces from agent trajectories to synthesize diverse training tasks and improves post-training performance through supervised fine-tuning.Agents and evaluation · Sep 2026 · Unverified1
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous DrivingQwen-Drive-1.0 is a vision-language foundation model for autonomous driving that unifies 3D perception, visual question answering, and motion planning via shared representations and staged training.Multimodal and robotics · Aug 2026 · Unverified3
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training StabilityQwen3.8-Flash-Next: a 125B mixture-of-experts model with 6B active that nearly matches its 397B predecessor at 1/9 the training compute.Architectures · Aug 20266
Topic
PaperCitations
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGemini 2.X model family, including Gemini 2.5 Pro and Flash, offers superior coding, reasoning, and multimodal understanding capabilities across a range of computational efficiencies.Google · Jul 2025 · Unverified4,134
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.DeepSeek · Dec 2025 · Unverified761
SkillOpt: Executive Strategy for Self-Evolving Agent SkillsSkillOpt introduces a systematic text-space optimizer for agent skills that trains skills as external agent state with stable updates and zero deployment inference overhead, achieving superior performance across multiple benchmarks and execution environments.Microsoft · May 2026 · Unverified85
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic ReasoningWe present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model.NVIDIA · Dec 2025 · Unverified81
Agent Learning via Early ExperienceEarly experience, using agent-generated interaction data without reward signals, improves policy effectiveness and generalization, serving as a bridge between imitation learning and reinforcement learning.Meta · Oct 2025 · Unverified64
In-the-Flow Agentic System Optimization for Effective Planning and Tool UseAgentFlow, a trainable agentic framework with in-the-flow optimization, enhances reasoning in large language models by coordinating specialized modules and outperforms top baselines across various tasks.Stanford University · Oct 2025 · Unverified58
About this paper
Authors
Rui Ye, Zhongwang Zhang, Kuan Li and 12 more
arXiv
2510.24699 · PDF
Venue
arXiv.org
Citations
77, 5 influential · Semantic Scholar
Upvotes
73 · Hugging Face
Lab
Alibaba (Qwen) · on Companies · on Quarterly

Changes

What changed
Influential citationsfirst count: 5Sep 25, 2026
Citationsfirst count: 77Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.