Details of Long-Sequence RL Training
nrehiew_ · x · 2026-07-09
The post summarizes the method's key changes: shifting from group-wise sampling to single-rollout, which necessitates reintroducing a value model to reduce variance.
It then lists several training details, including performing two gradient updates per batch for the value network, freezing attention to train only the MoE layers, computing GAE exclusively on action tokens, and implementing length-adaptive GAE.
Related event: GLM Team Proposes SAO Algorithm for Asynchronous Agent RL(15 posts)→
More from Research
- Structural ensembles beat single predictions in TCR:pMHC generalization study — quaidmorris · 2026-07-22
- Structural ensembles, not single predictions, drive robust TCR:pMHC generalization — quaidmorris · 2026-07-22
- A 3D ray plot shows how hard this Jacobian counterexample is to read — moultano · 2026-07-22
- LLM leaderboards are now often measuring the harness too, Gary Marcus warns — GaryMarcus · 2026-07-22
- New paper defines self-state attacks, showing OS defenses leave four agent-memory cases indistinguishable — Justgototheeffinmoon · 2026-07-22
- Krea 2 users recommend a two-pass Clownshark sampler setup for sharper image details — listopalafoto · 2026-07-22