Zhipu's SAO: single-rollout async RL trains stably for 1,000 steps, beats GRPO

teortaxesTex · x · 2026-08-21

Zhipu's paper SAO (Single-Rollout Asynchronous Optimization) tackles off-policy and stability challenges in asynchronous RL for agentic tasks:

SAO trains stably for 1,000 steps and consistently outperforms GRPO variants on agentic benchmarks like SWE-Bench Verified, BeyondAIME and IMOAnswerBench. It is already deployed in the agentic RL pipeline for GLM-5.2.

Original post →

More from Models

Models channel →