Extracting Robotic Action Signals from Egocentric Videos Using Only Open-Source Models

gui_penedo · x · 2026-08-07

Macrodata Labs released a new research blog exploring how to recover missing action signals (how hands move through 3D space) from egocentric videos using only open-source models. This is a crucial step for leveraging vast amounts of web video to train Vision-Language-Action (VLA) models.

Related event: Open-Source Pipeline Extracts Robot Actions from Egocentric Video(2 posts)→

Original post →

More from Embodied

Embodied channel →