Handling Hard Boundaries in Off-Policy Training
ziv_ravid · x · 2026-07-11
This segment covers an off-policy training approach: they discard the old policy and use the logprobs of the "current policy / rollout engine" as the importance ratio. Any samples exceeding the two-sided hard boundaries are directly masked out in the gradient, rather than applying standard clipping.
The core idea is that in agent training, they use stricter filtering to control off-policy errors instead of relying on mild clipping.
Related event: GLM Team Proposes SAO Algorithm for Asynchronous Agent RL(15 posts)→
More from Research
- Why a 1GW Chinese AI data center may be plausible after all — teortaxesTex · 2026-07-22
- Chinese AI labs are now treating distillation obfuscation as the top research topic — pmddomingos · 2026-07-22
- RSS launches under OMSF to push structural biology data modeling at scale — MoAlQuraishi · 2026-07-22
- enFoldX turns AlphaFold3 ensemble noise into a TCR–peptide–MHC predictor — quaidmorris · 2026-07-22
- enFoldX tops 8 neoantigen scans and an unseen-peptide benchmark — quaidmorris · 2026-07-22
- enFoldX reaches AUC 0.82 on human VDJdb and transfers to mouse at 0.76 — quaidmorris · 2026-07-22