Qwen Paper Proposes Elastic Horizon, a Closed-Loop Controller for Agentic RL Interaction Budgets

dair_ai · x · 2026-09-11

A Qwen team paper tackles interaction horizon scheduling in agentic RL: scaling per-episode environment interactions helps long-horizon agents, but existing curricula are open-loop and monotonic.

The authors define the "effective interaction frontier" — beyond it, extra interactions yield diminishing returns while costs grow linearly, with clear saturation plateaus on AppWorld and BFCL. Elastic Horizon is a closed-loop controller that tracks this boundary using the 90th percentile of successful trajectory lengths, a statistic already available during training.

Original post →

More from coding & agent

coding & agent channel →