JHU's Track, Articulate, Act reconstructs articulated objects and dexterous interactions from casual monocular videos
mangahomanga · x · 2026-09-17
Researchers at Johns Hopkins' Brains, Bots, and Behavior Lab introduce Track, Articulate, Act, a real-to-sim framework that reconstructs simulation-ready articulated objects (doors, drawers, cabinets, laptops, ovens) from a single casually captured monocular RGB video — no RGB-D, multi-view input, prior scans, manually specified joints, or robot demos needed.
Key insight: dense 3D point tracks are an embodiment-agnostic articulation cue — points on the fixed link stay stationary while points on moving links follow coherent revolute or prismatic motion. The pipeline repurposes pretrained models (SAM 3D for geometry, mesh segmentation, 3D scene flow), segments links, estimates joints and state trajectories, reconstructs the articulated asset, re-targets recovered 3D hand motion, and replays the interaction in MuJoCo physics simulation.
More from Embodied
- ModAR: a 30.1M robot world model trained from scratch beats a 6B video-model baseline — CSProfKGD · 2026-09-17
- Figure's Brett Adcock teases 'most important update' in company history, arriving tomorrow — adcock_brett · 2026-09-17
- MessyMem (CoRL 2026): persistent memory for robots via 3D scene graphs and VLM analysis — leto__jean · 2026-09-17
- Fortell founder spent six years building an AI hearing aid startup — RebeccaBellan · 2026-09-17
- Tesla Robotaxi resists passenger takeover attempts; commenters warn of lifetime bans — oilmutt · 2026-09-17
- Hours of zero-shot SO-101 testing fails to reproduce Astra VLA demo claims — chris_j_paxton · 2026-09-17