AC2 lets LLM RL train faster than GRPO by trusting critics on partial rollouts
stanfordnlp · x · 2026-10-06
Researchers propose Actor-Critic with Action Chunking (AC2), arguing GRPO-style RL needn't run every rollout to completion. A learned critic scores chunks of tokens, enabling training from partial rollouts — with parameter sharing between Q and policy and exploitation of easy LLM resets — reportedly training faster than GRPO. The work targets efficiency as billions flow into self-play and RSI, and has been amplified by Stanford AI researchers.
More from Research
- tszzl: Mechanistic interpretability is the bare minimum to make AI alignment an engineering discipline — tszzl · 2026-10-06
- Causal Decision-Making Preprint Presented at Simons Institute Trustworthy AI Workshop — murat_kocaoglu_ · 2026-10-06
- Claude produces O(n^1.9992) 3SUM algorithm with Lean proof, vetted by top experts — thegautamkamath · 2026-10-06
- ReSteer open-sourced: fixing VLA policies that ignore mid-execution instruction switches — siddkaramcheti · 2026-10-06
- Researcher presenting Latent Policy States in Reasoning Models at COLM this week — hunarbatra · 2026-10-06
- COLM 2026 Paper: Reasoning Fine-Tuning Induces Persistent Latent Policy States in LLMs — hunarbatra · 2026-10-06