First real-world RL fine-tuning on a VLA: policy turns jerky after ~10 rollouts

DominiqueCAPaul · x · 2026-09-28

The author shares a first real-world RL token experiment on a VLA, activating the learned residual policy only for the critical insertion step. Initial RLT rollouts work decently, but after 10 rollouts with updates the policy becomes visibly jerky, as shown in the video. Suspected causes: exploration noise injection or insufficient regularization of the residual policy—though noise should have shown up in the first rollouts too. Switching the RLT network off and letting the unedited VLA take over restores smooth motion and clean actuator insertion.

Original post →

More from Embodied

Embodied channel →