RL2-VLA: adaptive RL latent compositional steering with test-time scaling for VLA models

rsasaki0109 · x · 2026-08-18

Presents RL2-VLA, an adaptive reinforcement learning approach for Vision-Language-Action models. It relies on base VLA samples to approach the task object, then adaptively applies RL compositional steering when failure is preemptively detected, diversifying actions toward successful states, combined with test-time scaling.

Related event: RL²-VLA Boosts Robot Manipulation via Reinforcement Learning(2 posts)→

Original post →

More from Embodied

Embodied channel →