NVIDIA's Kimodo: Open-source text-to-3D-motion model runs on RTX 3090

maier_ak · x · 2026-08-31

NVIDIA released Kimodo, an open-source diffusion model that generates full-body 3D human or humanoid-robot motion from text prompts while honoring keyframe, foot-target, or path constraints. The 282M-parameter model achieves an average joint-position error of 3.2 cm, 71.9% R-precision, and FID 1.85 on the 700-hour Bones Rigplay set. It runs on a single RTX 3090 or even GPUs with <3 GB VRAM.

Related event: NVIDIA Open-Sources Kimodo for Text-to-3D Human Motion Generation(4 posts)→

Original post →

More from Embodied

Embodied channel →