RL for Ultra-Long Horizons: Shifting to Off-Policy and Critics

_AndrewZhao · x · 2026-08-21

As task horizons extend to 10+ hours or days, on-policy RL is no longer feasible. Andrew Zhao raises the question: how do we perform RL in this context? The discussion points towards off-policy, offline RL, and the use of critics.

Original post →

More from coding & agent

coding & agent channel →