Stanford RL method fixes VLA latency, lifting robot success from 42% to 97% with 10 minutes of data
burny_tech · x · 2026-09-21
A Stanford team (Chelsea Finn, Dorsa Sadigh et al.) published "Reinforcement Learning for Real-Time Vision-Language-Action Policies," tackling the core bottleneck of using large VLA models for reactive robot control: high inference latency means observations are stale by execution time, causing distribution shift and degraded reliability.
Approach
- Decouple slow, expressive action generation from fast, reactive editing: the large pretrained VLA proposes action chunks in the background, while a lightweight RL edit policy reacts to state changes using the latest observation.
- Built on the EXPO-FT RL fine-tuning framework, instantiated as Real-Time EXPO-FT — the first to bring RL fine-tuning up to real-time control requirements for dynamic real-world manipulation.
Results
- On the Kinetix benchmark, the delayed policy beat all non-delayed methods in 10/10 environments.
- On four dynamic real-world tasks (object passing, ball balancing, table soccer kicking, dynamic picking), with online robot data capped at 10 minutes, average success improved from 42% to 97%.
More from Embodied
- ZuckOff App Detects Nearby Meta Smart Glasses via Bluetooth Fingerprints, Tops 5,000 Downloads — LexiLove · 2026-09-21
- GaME (CVPR 2026): Gaussian mapping that forgets stale geometry as robot scenes change — lucacarlone1 · 2026-09-21
- Deadlifting 160kg at LUDUS: a striking strength demo — AIFlow_ML · 2026-09-21
- DeformSmith generates physics-grounded deformable assets for robot manipulation — Can Li · 2026-09-21
- Coowa breaks down COOWAM: on-device world model behind 10,000+ deployed city robots — 创业邦 · 2026-09-21
- Tesla's Texas Optimus line nears structural completion; Qiyuan claims a robot every 2.5 minutes — 创业邦 · 2026-09-21