QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents
Large language model agents tackling extreme-long horizon tasks generate massive interaction steps and tokens per rollout, causing severe GPU idling from execution variance and trajectory redundancy from non-linear branching. To address these issues, the authors present QwenGyre, an end-to-end online reinforcement learning framework. QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while a trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths. Scaled to Qwen~3.8 2.4T with 700K tokens per rollout, QwenGyre yields substantial absolute gains on NL2RepoBench and significant speedups over baseline methods.
QwenGyre yields a 6.0% absolute gain on NL2RepoBench (52.5% to 58.5%) in 48 steps.
QwenGyre delivers up to 1.85times speedups over Colocate.
QwenGyre delivers up to 1.78times speedups over Async.