FM-VLA Integrates Force Sensors to Resolve Ambiguity in Robot Policies

stepjamUK · x · 2026-08-06

Current mainstream Vision-Language-Action (VLA) models are fundamentally Markovian, mapping the current visual frame directly to the next action. To compensate for the lack of temporal context, the standard fix has been to add more vision: stacking history frames or extending image context.

However, for tasks like "pressing a button three times," the visual scene barely changes between presses, creating severe ambiguity for camera-based policies. FM-VLA makes a compelling case that force is the right channel to solve this. Wrist force sensors provide sharp, distinct spikes for each physical interaction, completely bypassing the limitations of visual ambiguity.

Original post →

More from Embodied

Embodied channel →