HarmoHOI generates multi-view hand-object videos and aligned 3D motion in one diffusion model
cn-scut · hf · 2026-07-21
HarmoHOI is a diffusion framework for synthesizing hand-object interaction videos across multiple views, while also producing globally aligned 3D point tracks.
What it does
- Jointly generates synchronized multi-view RGB motion videos and 3D motion trajectories.
- Treats point tracks as pseudo-videos to better align 3D geometry with 2D diffusion priors.
- Adds a global motion aligning module that refines coarse tracks into metric-scale trajectories.
Why it matters
- The paper argues that multi-view consistency needs both aligned 3D geometry and motion, not just stronger video priors.
- It uses a hybrid curriculum to transfer priors from single-view data to synchronized multi-view generation.
- Reported results claim state-of-the-art gains in visual quality, motion plausibility, and multi-view geometric consistency.
More from Embodied
- Openloong shows a wheeled humanoid robot autonomously hauling trash bins in Shanghai — CyberRobooo · 2026-07-21
- Open-AoE opens 2,000 hours of egocentric manipulation video for robot learning — inclusionAI · 2026-07-21
- Blender depth maps drive an LTX-2.3 IC-LoRA video workflow in ComfyUI — waterarttrkgl · 2026-07-21
- PROWL uses a world model to keep Minecraft agents exploring after failures — nathanbenaich · 2026-07-21
- LeRobot v0.6.0 adds end-to-end 3D depth training data for robots — RemiCadene · 2026-07-21
- NVIDIA brings its Cosmos 3 Edge world model to Jetson for on-device robot control — liu_mingyu · 2026-07-21