Stanford's AC2 beats GRPO with 2.5x fewer decoding FLOPs via action-chunked critic credit assignment

srush_nlp · x · 2026-10-02

A paper by Kaiyue Wen, Luke Bailey, Arvind Mahankali and Tengyu Ma introduces Actor-Critic with Action Chunking (AC2) for LLM RL training.

Problem: Standard RL algorithms credit every token with the same terminal-reward advantage; learned critics are usually trusted only as baselines, forcing every trajectory to roll out to completion.

Method:

Results: Training Qwen3-4B on FineProofs-RL with AC2 beats GRPO's peak validation score of 18.5% on IMO-ProofBench using 2.5x fewer decoding FLOPs, from 25% fewer training steps plus cheaper per-step rollouts.

Original post →

More from Models

Models channel →