JHU's Track, Articulate, Act reconstructs articulated objects from casual monocular video
mangahomanga · x · 2026-09-17
Johns Hopkins researchers (Jiaming Zhang, Homanga Bharadhwaj) present Track, Articulate, Act: from a single casually captured monocular RGB video — no RGB-D, multi-view, scans, or robot demos — the framework reconstructs simulation-ready articulated objects and hand-object interactions. Key insight: dense 3D point tracks are an embodiment-agnostic articulation cue, as points on the fixed link stay still while points on moving links follow coherent revolute/prismatic motion. It repurposes pretrained models for single-image 3D reconstruction, mesh segmentation, and 3D scene flow, and retargets recovered human interactions to an Allegro robot hand in MuJoCo.
More from Embodied
- Deep dive: what's really happening behind the Tesla Optimus, Figure and 1X NEO humanoid race — MickeySteamboat · 2026-09-17
- Figure teases an 'AI breakthrough' with a demo set for tomorrow — adcock_brett · 2026-09-17
- ModAR: a 30.1M robot world model trained from scratch beats a 6B video-model baseline — CSProfKGD · 2026-09-17
- Embodied AI is retracing the cobot path a decade later — yongqianme · 2026-09-17
- MessyMem (CoRL 2026): persistent memory for robots via 3D scene graphs and VLM analysis — leto__jean · 2026-09-17
- Fortell founder spent six years building an AI hearing aid startup — RebeccaBellan · 2026-09-17