VLA reinforcement learning updates are low-rank and concentrate in overlooked Timestep Modules
HOLILAB · hf · 2026-10-01
HOLILAB systematically studies how RL reshapes vision-language-action (VLA) policies in parameter space.
- Across flow-based VLAs (π₀.₅, GR00T N1.5/N1.6) on LIBERO, ManiSkill, MetaWorld, and CALVIN, RL induces substantially lower-rank updates, highly concentrated in the action expert's Timestep Modules — a small, previously overlooked component that captures a disproportionate share of RL's gains.
- RL specializes these modules to the discrete denoising timesteps used during rollouts; this discrete-timestep training underlies the low-rank updates.
- Among module outputs, the shift vector changes most distinctly under RL: shift update directions predict task success (ROC-AUC up to 99.6%), and their geometry correlates with cross-task transfer patterns.
- Steering along shift update directions further improves RL-trained policies without additional RL training.
Implications for more efficient, interpretable VLA post-training.
More from Embodied
- Personalized AI, personalized bugs: Dots users report wildly different flaws — ChrisGPT · 2026-10-01
- A 40-gram autonomous drone hunts mosquitoes by acoustic signature and sonar — TinfoilTricorn · 2026-10-01
- Google DeepMind Tokyo hiring senior research engineer for gaming AI agents — heiga_zen · 2026-10-01
- 5 robotics accounts that are '6 months ahead': Jim Fan, Physical Intelligence founders lead the list — RemiCadene · 2026-10-01
- DyRAD: radar novel view synthesis for dynamic driving scenes, recovering 90.7% vs 26.9% baseline — Technion · 2026-10-01
- NVIDIA's Ming-Yu Liu on open models, world models, and the future of physical AI — Practical AI · 2026-10-01