Skip to content

QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents research paper by Alibaba (Qwen), 2026

Alibaba (Qwen) · Sep 27, 2026 · Training and scaling · 43 upvotes · unverified 8 days ago

Read on arXiv

What it shows

Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout.

By Weiqi Wang, Yuxin Zhou, Mouxiang Chen and 9 more · arXiv 2609.33848 · PDF

UnverifiedHugging Face's summary; not yet checked by hand.

More from Alibaba (Qwen)

All 63
PaperCitations
Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief StatesLarge language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world.Agents and evaluation · Oct 2026 · Unverified4 days ago-
HappyWorld-BenchEvaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification.Agents and evaluation · Sep 2026 · Unverified2 weeks ago0
One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering AgentsRepository-level software engineering (SWE) comprises heterogeneous task categories, whose progress under pooled agentic reinforcement learning can be uneven: gains in some categories coincide with regressions in...Training and scaling · Sep 2026 · Unverified2 weeks ago0
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual DialogueWe define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model.Multimodal and robotics · Sep 2026 · Unverified2 weeks ago0
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use AgentsComputer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line.Agents and evaluation · Sep 2026 · Unverified2 weeks ago1
CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker DistillationCORE distills compositional ranking judgments from a cross-attentive reranker into an embedding model via synthesized multi-level candidates and a Rank-KL objective, improving compositional retrieval without degrading standard performance.Reasoning · Sep 2026 · Unverified4 weeks ago0
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal EnvironmentsTerminal-Universe reconstructs executable workspaces from agent trajectories to synthesize diverse training tasks and improves post-training performance through supervised fine-tuning.Agents and evaluation · Sep 2026 · Unverified4 weeks ago3
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous DrivingQwen-Drive-1.0 is a vision-language foundation model for autonomous driving that unifies 3D perception, visual question answering, and motion planning via shared representations and staged training.Multimodal and robotics · Aug 2026 · Unverified5 weeks ago7
Topic
PaperCitations
LoRA: Low-Rank Adaptation of Large Language ModelsLoRA fine-tunes a large model by training small low-rank matrices, cutting trainable parameters by 10,000 times.Microsoft · Jun 20215 years ago23.8k
Scaling Laws for Neural Language ModelsLanguage model loss falls as a smooth power law as model size, data and compute grow.OpenAI · Jan 20206 years ago9,190
Training Compute-Optimal Large Language ModelsChinchilla: for a fixed compute budget, train a smaller model on more data; parameters and tokens should grow together.Google DeepMind · Mar 20224 years ago3,835
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model ParallelismMegatron-LM splits each Transformer layer across GPUs to train models with billions of parameters.NVIDIA · Sep 20197 years ago3,236
DeepSeek LLM: Scaling Open-Source Language Models with LongtermismDeepSeek LLM, an open-source language model project, develops a large dataset and employs SFT and DPO to achieve performance surpassing LLaMA-2 70B and GPT-3.5 in various benchmarks and open-ended evaluations.DeepSeek · Jan 2024 · Unverified2 years ago861
LoRA Learns Less and Forgets LessLoRA learns less than full fine-tuning on code and math, but forgets less of what the model already knew.Databricks · May 20242 years ago422
About this paper
Authors
Weiqi Wang, Yuxin Zhou, Mouxiang Chen and 9 more
arXiv
2609.33848 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
43 · Hugging Face
Lab
Alibaba (Qwen) · on Companies · on Quarterly

Changes

What changed
New paperFound by the weekly scan, unverifiedOct 5, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.