UniMotion: Unified framework for motion, text, and vision achieves SOTA

jiqizhixin · x · 2026-08-23

UniMotion is a unified framework for motion-text-vision understanding and generation. It treats human motion as a continuous signal, pairing motion and images in a shared language model for seamless interaction. Using clever alignment tricks, it learns motion without images at test time and bootstraps with pure motion data. It achieves top scores across seven tasks, excelling at cross-modal combinations like text-to-motion generation and image editing via pose.

Original post →

More from Embodied

Embodied channel →