Tencent Hunyuan proposes SAT to stabilize asynchronous RL under stale rollouts
Tencent-Hunyuan · hf · 2026-07-22
Tencent Hunyuan presents SAT, a staleness-adaptive trust-region method for stabilizing asynchronous reinforcement learning.
Main contribution
- Uses detached sampled log-ratio as a practical proxy for rollout staleness.
- Identifies high-mismatch tails in each batch with staleness-based kernel scaling.
- Contracts only the sign-selected end of the PPO interval to preserve normal-token behavior while tightening risky updates.
Reported results
- Built on Qwen3-30B-A3B-Base with SGLang for inference and Megatron for training.
- On the asynchronous setup, SAT-GSPO w/ R3 reached the best observed AIME24 avg@8: 35.83 at lag 1 and 34.79 at lag 8.
- The authors argue that adaptive clipping plus routing replay jointly stabilize async RL under heterogeneous staleness.
Related event: Tencent Hunyuan Proposes SAT to Stabilize Asynchronous RL(2 posts)→
More from Research
- NSF and Astera Partner on Programmable Cloud Labs for AI-Ready Scientific Data — anshulkundaje · 2026-07-23
- Three classic improper integrals all collapse to √π or π — elonmusk · 2026-07-23
- Google Research studies AI agents for symptom interviews and diagnosis — gaganghotra_ · 2026-07-23
- Six practical ways to debug agent failures and keep performance from regressing — rdbms · 2026-07-23
- OSTP chief Michael Kratsios authors Genesis Mission report on AI for science — JungWooHa2 · 2026-07-23
- Hugging Face sandbox escape is being downplayed, says infosec researcher — proofreadre · 2026-07-23