Paper: Long-Horizon Agents Need Strong Pre-training and OPD
rohanpaul_ai · x · 2026-08-25
Addressing poor reliability in long-horizon agents, a new paper argues that post-training alone cannot fix weak foundations due to noisy trajectory error compounding. It proposes using clean world-models during pre-training and On-Policy Distillation (OPD) for post-training when rewards are sparse. Experiments show OPD handles long, noisy settings better than outcome-reward GRPO.
More from coding & agent
- Build RSS Reader and iOS Client with AI Vibe Coding — vista8 · 2026-08-25
- WSJ: Bosses push for public Slack channels to feed data to AI agents — rohanpaul_ai · 2026-08-25
- Claude Cowork retains build context better than standard chat — nptacek · 2026-08-25
- Qwen 27B Benchmark: Tools Lift Professional Finance Score to 98% — offgridai · 2026-08-25
- Heimdall: A CPU-Only Agent Memory System with Zero LLM Overhead — Slight-Parfait3679 · 2026-08-25
- Qwen 27B Coding Run Consumes 919M Input Tokens, Caching Cuts Costs — TheZachMueller · 2026-08-25