AC2 lets LLM RL train faster than GRPO by trusting critics on partial rollouts

stanfordnlp · x · 2026-10-06

Researchers propose Actor-Critic with Action Chunking (AC2), arguing GRPO-style RL needn't run every rollout to completion. A learned critic scores chunks of tokens, enabling training from partial rollouts — with parameter sharing between Q and policy and exploitation of easy LLM resets — reportedly training faster than GRPO. The work targets efficiency as billions flow into self-play and RSI, and has been amplified by Stanford AI researchers.

Original post →

More from Research

Research channel →