VideoMDM accepted to NeurIPS: 3D motion diffusion trained with only 2D supervision
orlitany · x · 2026-10-10
- VideoMDM, from Technion and NVIDIA researchers, is accepted to NeurIPS (Sydney). It trains 3D human motion diffusion models using only 2D poses from monocular video — no 3D ground truth required.
- Method: a pretrained 2D-to-3D lifter provides approximate 3D poses as a noisy teacher; the model denoises in 3D and is supervised purely in 2D via reprojection. Under mild assumptions, depth-weighted 2D reprojection loss equals direct 3D supervision in expectation; standard 3D regularizers are adapted to this setting.
- Results: FID 0.88 on HumanML3D vs 0.54 for fully 3D-supervised MDM; 2× lower MPJPE than WHAM on Fit3D; 64% human preference over MAS on NBA. Unlike inference-only lifting, it learns a coherent 3D motion manifold during training.
More from Embodied
- Agility Robotics lobbies at White House summit to ease humanoid deployment — xmercury_one · 2026-10-10
- Airbus Helicopters autonomy work touted as one of the coolest autonomy applications — BrettKrieger12 · 2026-10-10
- 15 EU robots showcased at Parliament, but no order rivals China's 8,500-unit State Grid plan — ingliguori · 2026-10-10
- Not all robots come with legs: Linus Ekenstam shares a claw-bot moment — LinusEkenstam · 2026-10-10
- Dev slams Vision Pro 2 UX: timeline mis-touches and unusable typing, urges Apple to build glasses — Kuprel · 2026-10-10
- Fully local AI plush toy: Whisper + Hermes 8B + Kokoro, zero cloud, MIT-licensed — msalsas · 2026-10-10