Researchers turn to DSRL to improve BC diffusion policies via latent-space RL
DominiqueCAPaul · x · 2026-09-03
Robotics researcher jannesdge is about to apply DSRL (Diffusion Steering via Reinforcement Learning) to improve behavior-cloned policies. From Sergey Levine et al.'s paper, DSRL runs RL over the latent-noise space of a diffusion policy, enabling sample-efficient autonomous real-world improvement with only black-box policy access. The catch: the policy must be steerable—producing visibly different trajectories under different noise initializations—and a quick test suggests this holds.
More from Embodied
- Agility Digit rearranges a full room of objects via RL teleop policy — chris_j_paxton · 2026-09-03
- RealSense handles robot 3D perception while QNX provides safety-certified real-time control for physical AI — pdamodaran · 2026-09-03
- Heart Aerospace's electric plane flew on just $5 of electricity — markjeffrey · 2026-09-03
- Uber lays off 3,300 people, 10% of staff, to flatten management and boost robotaxi bets — saibharadwaj · 2026-09-03
- Greg Madison's digital-twin house demo wows early testers — Scobleizer · 2026-09-03
- HT Robotics' ~$5k Mini Pi humanoid faces price pressure from $399 MiniDuck — philfung · 2026-09-03