Stanford's AC2 beats GRPO with 2.5x fewer decoding FLOPs via action-chunked critic credit assignment
srush_nlp · x · 2026-10-02
A paper by Kaiyue Wen, Luke Bailey, Arvind Mahankali and Tengyu Ma introduces Actor-Critic with Action Chunking (AC2) for LLM RL training.
Problem: Standard RL algorithms credit every token with the same terminal-reward advantage; learned critics are usually trusted only as baselines, forcing every trajectory to roll out to completion.
Method:
- AC2 assigns credit to action chunks (short continuations of trajectory prefixes), so the policy can update without observing a terminal reward
- "Local readiness" gates critic-based updates to problems where the critic is accurate enough
- The critic can be given a reference solution from a prior successful rollout
- Credit is scored over 10k-token chunks rather than individual tokens
Results: Training Qwen3-4B on FineProofs-RL with AC2 beats GRPO's peak validation score of 18.5% on IMO-ProofBench using 2.5x fewer decoding FLOPs, from 25% fewer training steps plus cheaper per-step rollouts.
More from Models
- Google clarifies Gemini TTS voice policy: custom voices only removed after one year of non-use — AI_Andrew · 2026-10-02
- Behavioural evals 'escaped containment' and cheated on their tests, says exec — gabriel1 · 2026-10-02
- KOL says Opus 5.5 is the first model to genuinely surprise him with its taste — Hesamation · 2026-10-02
- OpenAI clones TypeSafe's Jev just 14 days after debut, eroding its moat — JnBrymn · 2026-10-02
- Grok now answers questions directly inside X group chats via XChat — Baconbrix · 2026-10-02
- Daniel Bourke confirms SigLIP nuanced-text score example was real, from late 2024 — mrdbourke · 2026-10-02