QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents
Weiqi Wang, Yuxin Zhou, Mouxiang Chen, Siyuan Zhang, Yi Zhang, Yuyan Luo, Zhiyu Yin, Chencan Wu, Jiemin Jiang, Wentao Yao, Chujie Zheng, JianWei Zhang
cs.LG, cs.DC
2026-09-28
QwenGyre reallocates GPUs mid-rollout and caps branching trajectories. Qwen 3.8 2.4T goes 52.5% to 58.5% on NL2RepoBench in 48 steps, up to 1.78× faster than Async.
Issue-fixing benchmarks such as SWE-bench usually finish in tens of model-environment turns. Repository generation, full-system reimplementation, and codebase-wide migration stretch a single run to hours, hundreds of dependent interactions, and about 1M tokens across the lifecycle. The paper calls this regime XLong-horizon.
Online RL hits two walls at that scale. Colocate flips the entire GPU pool between rollout and training, so GPUs sit idle on the long tail. Async pins separate pools; waiting for a ready batch often outlasts a training step, and neither side can borrow spare cards. Deployed harnesses compact history, spawn sub-agents, and retry failed paths, turning one task into a branching graph with heavy shared prefixes.
The strongest harnesses are closed and still moving. Experiments here run Claude Code 2.1.220. Scheduling has to keep serving live executions; a dropped model call times out for infrastructure reasons, not policy ones.
QwenGyre splits three lifetimes: harness execution, GPU role, and training sample. Sandboxes and repos stay off the GPU workers. Model calls go through a proxy that reroutes them when a cell changes role.
Elastic scheduling carves GPUs into cells that share one training parallel layout. A cell can run rollout or host a full training replica. The scheduler tracks a waterlevel: dispatched executions that have not finished. When remaining rollout capacity still covers that work, cells move to training in a fixed order. The core cell owns the authoritative weights and the optimizer. Satellites pull a parameter snapshot over RDMA via Mooncake and join the in-flight batch. Streaming training hands out micro-steps; idle cells claim more work.
During a role change the proxy retries in-flight requests on other rollout engines and can migrate KV cache over RDMA, then swaps weights. Measured switch times are 8.52s rollout to train and 3.46s the other way, small next to hour-long runs.
The trajectory processor records exact tokens and behavior log-probs with TITO (token-in, token-out) and stores them in a prefix-sharing tree. After timeout, assessable partial artifacts still get a task score instead of a blanket failure. Evaluator failure is not a valid zero and is dropped. Each execution keeps at most Jmax=5 trajectories, ranked main > main-summary > sub-agent > sub-agent-summary. Shared targets count once. Loss averages over trainable tokens inside an execution, then over executions in the batch, so extra paths do not overweight a task. The optimizer is a GSPO variant: token importance capped at 5, sequence clip [0.995, 1.005], 16 rollouts per group.
Training uses Megatron; rollout uses SGLang. Baselines are Colocate and Async at equal GPU budget and matched mean scheduling staleness E[d], the expected gap between the policy version at dispatch and the version just before the consuming step.
| Setup | Baseline | Result |
| Qwen 3.6 122B, NL2RepoBench, 48 steps | Async | 1.42-1.53× cumulative |
| Same | Colocate | 1.36-1.47× |
| Qwen 3.8 2.4T, NL2RepoBench, 48 steps | Async 134.55h | 75.42h, 1.78× |
| Same | Colocate 91.47h | 1.21× |
| 122B DeepSWE, 24 steps | Async / Colocate | up to 1.57× / 1.82× |
| 122B TerminalBench, 24 steps | Async / Colocate | up to 1.43× / 1.85× |
On Qwen 3.8 2.4T the eval passrate moves from 52.48% to 58.54%. Training-score curves track both baselines, so the six points come from running 48 RL steps on this stack. The scheduler does not pull extra score over Colocate.
On 122B NL2RepoBench, a rollout averages 1.93h and a query 2.96h; 9.51% of queries take at least four hours. Harness timeout is 6.25%, overall timeout 0.87%. The 2.4T runs raise the timeout budget and still see 61.1% of queries at or above four hours; in 38.2% of those, the longest rollout dies on overall timeout. When runs pile up against the limit, cells rarely free early, and the Colocate gap shrinks to 1.21×. Async's pinned training pool sits idle through those long executions; giving that capacity to rollout is where the 1.78× comes from.
Ablations (122B, first 12 steps, E[d]=1.5): full config E0 takes 16.12h. Streaming Async (A3) still takes 22.18h; Colocate plus standalone and streaming (C2) takes 19.23h. Fine-grained cells add another 1.19× over C2. Single-step bursts (E1) cost 10.9% versus E0 and still beat streaming Async by 1.24×. Merging eight 4-node cells into four 8-node cells slows E0 by 2.2%. Sampling only the main trajectory hurts scores and inflates gradient norms; Jmax=5 matches the uncapped score at 74.8% of uncapped forward-backward time, a 25.2% cut.
This is systems work for hour-scale, near-million-token agent RL behind a black-box harness, not a new objective. It pays off if you already run Claude Code-like harnesses on repo-scale tasks and can slice tens to hundreds of GPUs into cells.
TideRL, BiDiRL, DynaResize, and Libra also move the rollout/train boundary. QwenGyre keeps live executions alive across role changes and bounds branching traces by role, training shared prefixes once. For a team training a Qwen-scale coding agent, 48 steps for about six NL2Repo points and wall time from 134h down to 75h is a concrete trade. The paper does not say whether the framework is open-sourced.
The authors flag two constraints. Elastic scheduling needs several independently switchable cells and enough leftover rollout capacity for unfinished work; a tiny GPU budget cannot host that topology, and smaller budgets were not tested. With more than one training step per burst, streaming assigns ready groups to successive updates in arrival order, so a burst cannot globally shuffle minibatches. The effect of that order on learning quality is unmeasured.
The 2.4T six-point lift uses internal training data and a Qwen 3.7-Max judge at train time; official eval follows the NL2RepoBench protocol, a different ruler. High timeout rates shrink the Colocate gap: if everyone dies on the deadline, there is little idle capacity to reclaim.