NUS's RL² Framework Boosts VLA Out-of-Domain Success Rates by 17%
NationalUniversityofSingapore · hf · 2026-08-03
National University of Singapore introduced RL², an adaptive inference-time steering framework that trains a lightweight offline RL policy on the latents of Vision-Language-Action (VLA) models.
- Mechanism: It composes the flow velocity of an RL policy (trained on expressive latents from the VLA action expert) with the frozen VLA during inference. The steering is activated adaptively—only when the base VLA is predicted to fail—preventing unnecessary perturbation of already accurate actions.
- Results: Across SIMPLER and PolaRiS benchmarks, RL² improves out-of-domain success rates by up to 17.3%. Real-world experiments confirm these simulation gains successfully transfer to physical robotic deployments.
More from Embodied
- Buenos Aires robotics meetup to feature ESP32 hands-on workshop and machine economy talk — StewartalsopIII · 2026-08-03
- Open-Source System Serves VLA Models to 10+ Robots on a Single GPU — danfei_xu · 2026-08-03
- Open-Source Quadruped Robot Framework OpenCat Hits 5k Stars — tom_doerr · 2026-08-03
- Qiuzhi Tech Raises Over $200M in Two Months, Unveils Embodied AI Model ORION-0 — 创业邦 · 2026-08-03
- Local DGX Spark Runs Minimax H3: 720p Video in 1h44m — Saren-WTAKO · 2026-08-03
- OpenRoboto Robotics Challenge: Fine-tune π0.5 to Win Token Rewards — const_reborn · 2026-08-03