Driving · 2026-09-30
QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents
Weiqi Wang, Yuxin Zhou, Mouxiang Chen, Siyuan Zhang, Yi Zhang, Yuyan Luo, Zhiyu Yin, Chencan Wu, Jiemin Jiang, Wentao Yao, Chujie Zheng, JianWei Zhang
Auto-summarized: this summary was generated by a language model from the paper’s text and has not been reviewed by an editor. Check the paper before relying on it. Benchmark numbers appear only after a human has verified them.
TL;DR
QwenGyre is an end-to-end framework for x Long-horizon online RL with black-box LLM agent harnesses. It uses an elastic scheduler to elastically reallocate GPUs between rollout and training without interrupting live executions, and a trajectory processor that reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training cost. On NL2RepoBench with Qwen3.8 2.4T, it improves score in 48 steps and reports end-to-end speedups over Colocate and Async.
Why it matters
x Long-horizon agentic RL has two bottlenecks the authors target: (1) severe execution variance and rollout delays that leave GPUs idle, and (2) non-linear trajectory branching that creates redundant training data and high overhead. QwenGyre’s elastic GPU reallocation and trajectory deduplication aim to improve both training efficiency and model performance in this setting.
Method
- Elastic scheduler: tracks rollout waterlevel and reallocates GPUs between rollout and centralized data-parallel training at burst boundaries while preserving ongoing harness executions.
- Trajectory processor: records model calls into trajectory trees with shared prefixes, evaluates preserved workspace including partial progress after timeout, and selects a bounded number of trajectories per execution by role priority.
- Training data construction: computes group-relative advantages from original execution rewards; uses on-tree masking so shared targets contribute once; averages token losses within each execution to prevent executions with more paths from overweighting losses.
Abstract (from arXiv)
Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental challenges: (1) severe execution variance and prolonged rollout delays cause massive GPU idling; and (2) complex non-linear branching generates massive trajectory redundancy, crippling training efficiency. To address these, we presents QwenGyre, an end-to-end framework for xlong-horizon online RL. QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs. Scaled to our flagship model, Qwen~3.8 2.4T, with 700K tokens per rollout, QwenGyre yields a 6.0% absolute gain on NL2RepoBench (52.5% $\to$ 58.5%) in 48 steps. Across our evaluations on diverse domains of training datasets, QwenGyre delivers up to $1.85\times$ and $1.78\times$ speedups over Colocate and Async, respectively.
Related papers
- TD-MPC2: Scalable, Robust World Models for Continuous Control
- FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales
- Transformers are Sample-Efficient World Models
- Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks
- Mastering Atari with Discrete World Models