SAO Outperforms GRPO in Coding and Reasoning Benchmarks

Researchers introduced SAO, an asynchronous RL algorithm designed for long-horizon agent tasks. SAO outperforms GRPO and its variants across multiple coding and reasoning benchmarks, with performance gains primarily attributed to Direct Double-Sided Importance Sampling (DIS).

2026-07-14 ~ 2026-07-15 · 3 related posts