JHU's Track, Articulate, Act reconstructs articulated objects and dexterous interactions from casual monocular videos

mangahomanga · x · 2026-09-17

Researchers at Johns Hopkins' Brains, Bots, and Behavior Lab introduce Track, Articulate, Act, a real-to-sim framework that reconstructs simulation-ready articulated objects (doors, drawers, cabinets, laptops, ovens) from a single casually captured monocular RGB video — no RGB-D, multi-view input, prior scans, manually specified joints, or robot demos needed.

Key insight: dense 3D point tracks are an embodiment-agnostic articulation cue — points on the fixed link stay stationary while points on moving links follow coherent revolute or prismatic motion. The pipeline repurposes pretrained models (SAM 3D for geometry, mesh segmentation, 3D scene flow), segments links, estimates joints and state trajectories, reconstructs the articulated asset, re-targets recovered 3D hand motion, and replays the interaction in MuJoCo physics simulation.

Related event: JHU's Track, Articulate, Act Rebuilds Articulated Objects from Monocular Video for Robots(5 posts)→

Original post →

More from Embodied

Embodied channel →