Qwen Paper Proposes Elastic Horizon, a Closed-Loop Controller for Agentic RL Interaction Budgets
dair_ai · x · 2026-09-11
A Qwen team paper tackles interaction horizon scheduling in agentic RL: scaling per-episode environment interactions helps long-horizon agents, but existing curricula are open-loop and monotonic.
The authors define the "effective interaction frontier" — beyond it, extra interactions yield diminishing returns while costs grow linearly, with clear saturation plateaus on AppWorld and BFCL. Elastic Horizon is a closed-loop controller that tracks this boundary using the 90th percentile of successful trajectory lengths, a statistic already available during training.
More from coding & agent
- Astra builds a surprisingly polished Catan game in three.js, reusing past UI and 3D assets — FinanceYF5 · 2026-09-11
- Open-Source Tool Highlights the Exact PDF Paragraphs Behind AI Answers — Flat-Phone-1596 · 2026-09-11
- OpenAI Codex may issue another usage reset this weekend, says Codex lead resets happen — umesh_ai · 2026-09-11
- Dev swaps gemini-3.8 for gemini-3.5-flash-lite in his MCP harness at a fraction of cost — julianharris · 2026-09-11
- ComfyUI Style Explorer Adds LoRA Preview Catalog and Sharing — neonsparksuk · 2026-09-11
- User burns $200 of Codex credits in one agent turn — 4,700 of 5,000 credits, task unfinished — RileyRalmuto · 2026-09-11