AC2: Action-Chunked Actor-Critic Trains Faster Than GRPO with Less Compute

Stanford researchers introduced AC2 (Actor-Critic with Action Chunking), which trusts the critic to evaluate token chunks and train on partial rollouts. It cuts decoding compute by roughly 2.5x while outperforming GRPO.

2026-10-02 ~ 2026-10-03 · 2 related posts