AC2: Action-Chunked Actor-Critic Trains Faster Than GRPO with Less Compute
Stanford researchers introduced AC2 (Actor-Critic with Action Chunking), which trusts the critic to evaluate token chunks and train on partial rollouts. It cuts decoding compute by roughly 2.5x while outperforming GRPO.
2026-10-02 ~ 2026-10-03 · 2 related posts
- Stanford's AC2 beats GRPO with 2.5x fewer decoding FLOPs via action-chunked critic credit assignment — srush_nlp · 2026-10-02
- AC2: actor-critic with action chunking enables partial rollouts, trains faster than GRPO — ZeYanjie · 2026-10-03