Zhipu's SAO: single-rollout async RL trains stably for 1,000 steps, beats GRPO
teortaxesTex · x · 2026-08-21
Zhipu's paper SAO (Single-Rollout Asynchronous Optimization) tackles off-policy and stability challenges in asynchronous RL for agentic tasks:
- Replaces GRPO-style group sampling with single-rollout sampling (one rollout per prompt)
- Adds practical value-model training designs
- Introduces a strict two-sided token-level clipping strategy for stability
SAO trains stably for 1,000 steps and consistently outperforms GRPO variants on agentic benchmarks like SWE-Bench Verified, BeyondAIME and IMOAnswerBench. It is already deployed in the agentic RL pipeline for GLM-5.2.
More from Models
- Reddit Tries Recursive Meta-Prompting to Cut Qwen's Thinking Tokens by Half — Rare_Potential_1323 · 2026-08-21
- NVIDIA Details Qwen3.8-2.4T Deployment on GB300, Achieving >4K Tokens/s per GPU — PyTorch · 2026-08-21
- Tutorial: Training a local LLM on a new domain via continued pretraining — funJS · 2026-08-21
- Qwen 3.8 27B two-shots a playable 3D amusement park game in the browser — RandumbRedditor1000 · 2026-08-21
- Gemini's Safety Filters Too Strict? Rejects Kissing and Roadside Photos — Dry-Sympathy-3182 · 2026-08-21
- 0.63M Parameter Verifier Matches 7B Models in Specific Tasks — jm_alexia · 2026-08-21