NVIDIA's Kimodo Turns Text Into Full-Body 3D Human and Robot Motion

maier_ak · x · 2026-08-31

A write-up on NVIDIA's Kimodo, the first open-source diffusion model that generates full-body 3-D human and humanoid-robot motion from natural-language prompts while honoring keyframe, foot-target and path constraints, running on a single RTX 3090 or under 3 GB VRAM.

Related event: NVIDIA Open-Sources Kimodo for Text-to-3D Human Motion Generation(4 posts)→

Original post →

More from Embodied

Embodied channel →