Researchers turn to DSRL to improve BC diffusion policies via latent-space RL

DominiqueCAPaul · x · 2026-09-03

Robotics researcher jannesdge is about to apply DSRL (Diffusion Steering via Reinforcement Learning) to improve behavior-cloned policies. From Sergey Levine et al.'s paper, DSRL runs RL over the latent-noise space of a diffusion policy, enabling sample-efficient autonomous real-world improvement with only black-box policy access. The catch: the policy must be steerable—producing visibly different trajectories under different noise initializations—and a quick test suggests this holds.

Original post →

More from Embodied

Embodied channel →