Handling Hard Boundaries in Off-Policy Training
ziv_ravid · x · 2026-07-11
This segment covers an off-policy training approach: they discard the old policy and use the logprobs of the "current policy / rollout engine" as the importance ratio. Any samples exceeding the two-sided hard boundaries are directly masked out in the gradient, rather than applying standard clipping.
The core idea is that in agent training, they use stricter filtering to control off-policy errors instead of relying on mild clipping.
Related event: GLM Team Proposes SAO Algorithm for Asynchronous Agent RL(15 posts)→
More from Research
- Michael Levin publishes peer-reviewed Platonic Space paper, his most controversial yet — drmichaellevin · 2026-09-11
- GameWorld wins Best Paper Runner-Up at ECCV 2026 Multimodal Digital Agents Workshop — MikeShou1 · 2026-09-11
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Sample selection and ordering matter a lot in LLM training: DataFlex makes data scheduling dynamic — Puzzleheaded_Box2842 · 2026-09-11
- Jeff Heaton's Intro to the Math of Neural Networks eBook Is Free to Download — blaizedsouza · 2026-09-11
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11