RL²-VLA Boosts Robot Manipulation via Reinforcement Learning
RL²-VLA is an adaptive inference-time steering framework for vision-language-action models. It trains a lightweight offline RL flow-matching policy that adaptively intervenes in the VLA's latent space when failure is anticipated, improving robot manipulation success rates.
2026-08-18 ~ 2026-08-18 · 2 related posts
- RL2-VLA: adaptive RL latent compositional steering with test-time scaling for VLA models — rsasaki0109 · 2026-08-18
- RL²-VLA boosts robot manipulation via RL steering — rsasaki0109 · 2026-08-18