Stanford's Real-Time EXPO-FT Brings RL to Real-Time VLA Robot Policies, Beats RTC

StanfordAILab · x · 2026-09-24

Stanford researchers (Perry Dong, Kuo-Han Hung, Dorsa Sadigh, Chelsea Finn) introduce Real-Time EXPO-FT, an RL fine-tuning framework for real-time vision-language-action policies. It tackles latency-induced distribution shift in dynamic tasks via a three-stage design: a large VLA base slowly proposes action chunks, a small edit policy revises them using observations at execution time, and a Q-function selects the best candidate on the fly. It unlocks π0.5 on tasks like ball balancing and striking, significantly outperforming RTC baselines. Paper and code are public.

Related event: Stanford's Real-Time EXPO-FT Makes VLA Robots Real-Time(2 posts)→

Original post →

More from Embodied

Embodied channel →