RL for Ultra-Long Horizons: Shifting to Off-Policy and Critics
_AndrewZhao · x · 2026-08-21
As task horizons extend to 10+ hours or days, on-policy RL is no longer feasible. Andrew Zhao raises the question: how do we perform RL in this context? The discussion points towards off-policy, offline RL, and the use of critics.
More from coding & agent
- Investigating Overhead and Decision Fatigue in Managing Agents at Scale — zakelfassi · 2026-08-21
- Workflow Tip: Connecting ChatGPT Pro to GitHub Beats Using Codex Alone — jdjohnson · 2026-08-21
- AMD MI300x vs NVIDIA H100: Real-world agent coding benchmark — locker73 · 2026-08-21
- CLI coding agent fx update: 6MB binary, WASM support, instant startup — evilrabbit_ · 2026-08-21
- Rebuttal: AI coding is iteration, not recursion — gerardsans · 2026-08-21
- Huzzah editor lets you write pseudocode while AI fills in the rest — gregbarbosa · 2026-08-21