Action Chunks Boost Contrastive RL by 31.7% Offline and 93.1% Online
ben_eysenbach · x · 2026-09-03
The arXiv paper Three Steps at a Time: Learning Representations from Action Sequences in Contrastive RL extends contrastive RL (CRL) from single-step actions to action chunks, yielding +31.7% across 18 offline and +93.1% across 11 online benchmark environments.
Action-chunking gains are usually attributed to modeling non-Markovian temporally extended policies and propagating unbiased multi-step returns. The authors find these arguments only partially apply to CRL: empirically, an action chunk carries more information about the goal than a single action, measurably improving the critic's representations and making the algorithm significantly more effective.
Paper, website and code are all publicly available.
More from Research
- Dev scrapes 5.94B TikTok videos and 3.23B profiles in 3 weeks, uploads dataset to Hugging Face — DataShack · 2026-09-03
- New hypothesis paper argues consciousness and cognition are separate evolutionary lineages, with six falsifiable predictions — rjhaier · 2026-09-03
- Ask HN-style: where to find legally usable datasets for advanced chord recognition? — DiscoramaMusic · 2026-09-03
- Revera: Lean-verified POSIX regex engine with identical output in 6 languages — jedisct1 · 2026-09-03
- Yale trio's new paper: mechanism design for AI agents with unknown alignment — Afinetheorem · 2026-09-03
- PufferLib 5.0 Self-Play Trains 10-Ship Duel in 5 Minutes on a $700 PC — yacineMTB · 2026-09-03