RL²-VLA boosts robot manipulation via RL steering

rsasaki0109 · x · 2026-08-18

RL²-VLA is an adaptive inference-time steering framework that improves robotic manipulation by applying Reinforcement Learning on VLA latents. It trains a lightweight offline RL flow-matching policy and steers the base VLA during inference. The method activates compositional steering mainly when failure is predicted. Benchmarks show average success rate improvements of 10.1% on SIMPLER, 8.9% on PolaRiS, and 19.5% on real robots.

Related event: RL²-VLA Boosts Robot Manipulation via Reinforcement Learning(2 posts)→

Original post →

More from Embodied

Embodied channel →