Tencent Hunyuan proposes SAT to stabilize asynchronous RL under stale rollouts
Tencent-Hunyuan · hf · 2026-07-22
Tencent Hunyuan presents SAT, a staleness-adaptive trust-region method for stabilizing asynchronous reinforcement learning.
Main contribution
- Uses detached sampled log-ratio as a practical proxy for rollout staleness.
- Identifies high-mismatch tails in each batch with staleness-based kernel scaling.
- Contracts only the sign-selected end of the PPO interval to preserve normal-token behavior while tightening risky updates.
Reported results
- Built on Qwen3-30B-A3B-Base with SGLang for inference and Megatron for training.
- On the asynchronous setup, SAT-GSPO w/ R3 reached the best observed AIME24 avg@8: 35.83 at lag 1 and 34.79 at lag 8.
- The authors argue that adaptive clipping plus routing replay jointly stabilize async RL under heterogeneous staleness.
Related event: Tencent Hunyuan Proposes SAT to Stabilize Asynchronous RL(2 posts)→
More from Research
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11
- Catholic University of Chile researcher: scaling AI feedback is key to sustainable medical education — julianvarascom · 2026-09-11
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11