JHU's Track, Articulate, Act reconstructs articulated objects from casual monocular video

mangahomanga · x · 2026-09-17

Johns Hopkins researchers (Jiaming Zhang, Homanga Bharadhwaj) present Track, Articulate, Act: from a single casually captured monocular RGB video — no RGB-D, multi-view, scans, or robot demos — the framework reconstructs simulation-ready articulated objects and hand-object interactions. Key insight: dense 3D point tracks are an embodiment-agnostic articulation cue, as points on the fixed link stay still while points on moving links follow coherent revolute/prismatic motion. It repurposes pretrained models for single-image 3D reconstruction, mesh segmentation, and 3D scene flow, and retargets recovered human interactions to an Allegro robot hand in MuJoCo.

Related event: JHU's Track, Articulate, Act Rebuilds Articulated Objects from Monocular Video for Robots(5 posts)→

Original post →

More from Embodied

Embodied channel →