First real-world RL fine-tuning on a VLA: policy turns jerky after ~10 rollouts
DominiqueCAPaul · x · 2026-09-28
The author shares a first real-world RL token experiment on a VLA, activating the learned residual policy only for the critical insertion step. Initial RLT rollouts work decently, but after 10 rollouts with updates the policy becomes visibly jerky, as shown in the video. Suspected causes: exploration noise injection or insufficient regularization of the residual policy—though noise should have shown up in the first rollouts too. Switching the RLT network off and letting the unedited VLA take over restores smooth motion and clean actuator insertion.
More from Embodied
- HSImul3R (ECCV 2026): simulation-ready human-scene reconstruction judged by physics, not looks — jiqizhixin · 2026-09-28
- World's largest humanoid robot livestream performance draws crowd of 10,000+ — kernelangus420 · 2026-09-28
- One-arm VR intervention during bimanual DAgger feels like an AI brain chip — neurosp1ke · 2026-09-28
- USTC's VLA-Precision: online RL for VLAs hits 98.3% success on precision chemistry tasks — ustc · 2026-09-28
- Robot training fix: cutting noise injection from 2.5% to 0.25% of joint range removes jerkiness — DominiqueCAPaul · 2026-09-28
- AHa-3D turns ordinary indoor videos into editable Blender 3D scenes — ducha_aiki · 2026-09-28