Qwen Open-Sources QwenGyre RL Framework for xLong-Horizon Agent Training

Qwen · hf · 2026-09-29

Qwen introduces QwenGyre, an end-to-end online RL framework for extremely long-horizon agents whose rollouts span hours and 1M tokens. It elastically reallocates GPUs between rollout and training without interrupting executions, and reconstructs/scores/deduplicates branching trajectories. Training Qwen3.8 2.4T with 700K-token rollouts yields a 6.0% absolute gain on NL2RepoBench (52.5%→58.5%) in 48 steps, with up to 1.85x/1.78x speedups over Colocate/Async baselines.

Original post →

More from coding & agent

coding & agent channel →