Action Chunks Boost Contrastive RL by 31.7% Offline and 93.1% Online

ben_eysenbach · x · 2026-09-03

The arXiv paper Three Steps at a Time: Learning Representations from Action Sequences in Contrastive RL extends contrastive RL (CRL) from single-step actions to action chunks, yielding +31.7% across 18 offline and +93.1% across 11 online benchmark environments.

Action-chunking gains are usually attributed to modeling non-Markovian temporally extended policies and propagating unbiased multi-step returns. The authors find these arguments only partially apply to CRL: empirically, an action chunk carries more information about the goal than a single action, measurably improving the critic's representations and making the algorithm significantly more effective.

Paper, website and code are all publicly available.

Original post →

More from Research

Research channel →