Video Generation Models for 4D Hand Motion Capture

新智元 · wechat · 2026-07-13

A new method called ACE-ViDiHand was introduced, applying video generation models directly to 4D hand motion capture, bypassing the traditional pipeline of "first detect hand, then estimate pose." The author claims it performs strongly on three benchmarks: ARCTIC, HOT3D, and HOI4D, significantly outperforming previous methods, especially in occlusion scenarios.

The core idea: first, overlay hand renderings onto video frames, forcing the model to internally maintain the 3D hand state during occlusion; then, directly read hand poses from the intermediate layer features of the DiT. The article stresses it requires no detectors, cropping, motion completion, or test-time optimization, outputting a dual-hand 4D trajectory in a single forward pass from full-frame video.

The article further elevates this to a data gateway for embodied AI: if hand movements can be stably extracted from massive first-person videos, "in-the-wild videos" can be transformed into learnable data assets. The author believes this represents a paradigm shift in 4D hand reconstruction from ad-hoc patch solutions toward "inheriting video world model capabilities."

Original post →

More from Embodied

Embodied channel →