UniMotion: Unified framework for motion, text, and vision achieves SOTA
jiqizhixin · x · 2026-08-23
UniMotion is a unified framework for motion-text-vision understanding and generation. It treats human motion as a continuous signal, pairing motion and images in a shared language model for seamless interaction. Using clever alignment tricks, it learns motion without images at test time and bootstraps with pure motion data. It achieves top scores across seven tasks, excelling at cross-modal combinations like text-to-motion generation and image editing via pose.
More from Embodied
- Chinese 'Lightning' Humanoid Runs 100m in 9.32s, Nearing Bolt's Record — minchoi · 2026-08-23
- Robots unlock new skill: capable of climbing steep stairs — Dr_Singularity · 2026-08-23
- RTX 5090 Benchmarks Minimax H3: 12 Seconds for 20 Steps — Grinderius · 2026-08-23
- China's AI robot hits 14.5 m/s, expert comment on design philosophy — kevinakwok · 2026-08-23
- Robot VR teleop demo: 4ms latency, 60Hz for spreading Nutella — neurosp1ke · 2026-08-23
- WHRG’26 Features First Live Human-Robot Doubles Tennis Match with Galbot — Distinct-Question-16 · 2026-08-23