MoSE3 estimates per-pixel SE(3) transforms from video, NeurIPS Spotlight
du_yilun · x · 2026-10-06
MoSE3 has been accepted as a NeurIPS 2026 Spotlight. Key technical points:
- Estimates an SE(3) transform at every pixel from monocular video, in a shared world coordinate frame and in a single forward pass
- Captures rigid, articulated, and deformable motion
- The reposter highlights its potential to convert generative video model outputs into executable 3D robot actions, bridging video generation and robotic manipulation
More from Embodied
- ProgressCompass: context injection cuts embodied progress reward model error by 63% — Jianshu Zhang · 2026-10-07
- Video2Skill benchmark: most of 19 open-source VLMs fail at streaming embodied skill discovery — Jianshu Zhang · 2026-10-07
- MM-ABC robot foundation model hits 83% on real-world mobile manipulation tasks — Qiwei Liang · 2026-10-07
- Cortex Harness lets frontier models run real robots, with human takeovers feeding training data — kimmonismus · 2026-10-07
- 60 builders to voice-enable a Furby at SF Tech Week hardware hackathon — AssemblyAI · 2026-10-07
- Browser-based Bluetooth support lets you teleoperate robots with a PS gamepad — chrismatthieu · 2026-10-07