QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents research paper by Alibaba (Qwen), 2026
Alibaba (Qwen) · Sep 27, 2026 · Training and scaling · 43 upvotes · unverified 8 days ago
What it shows
Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout.
By Weiqi Wang, Yuxin Zhou, Mouxiang Chen and 9 more · arXiv 2609.33848 · PDF
UnverifiedHugging Face's summary; not yet checked by hand.
More from Alibaba (Qwen)
All 63Other training and scaling papers
TopicAbout this paper
- Authors
- Weiqi Wang, Yuxin Zhou, Mouxiang Chen and 9 more
- arXiv
- 2609.33848 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 43 · Hugging Face
- Lab
- Alibaba (Qwen) · on Companies · on Quarterly
Changes
| What changed | |
|---|---|
| Oct 5, 2026today | New paperFound by the weekly scan, unverifiedOct 5, 2026today |