Qwen Open-Sources QwenGyre RL Framework for xLong-Horizon Agent Training
Qwen · hf · 2026-09-29
Qwen introduces QwenGyre, an end-to-end online RL framework for extremely long-horizon agents whose rollouts span hours and 1M tokens. It elastically reallocates GPUs between rollout and training without interrupting executions, and reconstructs/scores/deduplicates branching trajectories. Training Qwen3.8 2.4T with 700K-token rollouts yields a 6.0% absolute gain on NL2RepoBench (52.5%→58.5%) in 48 steps, with up to 1.85x/1.78x speedups over Colocate/Async baselines.
More from coding & agent
- KernelZero-7B co-evolution beats Claude 4.5 Sonnet on CUDA kernel generation — Changxin Ke · 2026-09-29
- AgentHop: an MCP server for E2E-encrypted agent-to-agent chat via one-time pairing codes — Inevitable-Back620 · 2026-09-29
- Triage agent design: 1-2 high-signal questions rescued B2B reps drowning in 60% low-intent chats — hubtyper · 2026-09-29
- Separating outcome memories from return memories in the Hindsight agent memory system — Harishkumar79 · 2026-09-29
- Noah Shunn, 23: Reflexion Author Who Beat GPT-4 on HumanEval, Now Agent Pioneer at Sierra — vaibhavbetter · 2026-09-29
- OpenAI Codex CLI 0.159.0 ships instant interrupt, Mermaid rendering and Windows fixes — github-actions[bot] · 2026-09-29