UMO: unified in-context framework turns text2motion model into multi-task motion foundation
Michael_J_Black · x · 2026-09-11
UMO (Brown, MIT, Max Planck, Meta Reality Lab, HKU; ECCV 2026) builds on a pre-trained 3D text-to-motion model and uses a unified in-context learning framework with three meta-operations to handle diverse motion tasks — in-domain text-to-motion generation plus out-of-domain temporal tasks like keyframe infilling, prediction, backcasting and in-betweening — beating task-specific baselines. Paper and code are public.
More from Embodied
- Reachy Mini robot gains spotlight after joining NVIDIA's portfolio — RachelVT42 · 2026-09-11
- NTU spin-off Ropedia raises $22M to train robots from wearable first-person data — liuziwei7 · 2026-09-11
- Why robotics folks stay quiet on AGI: dishes and laundry first, hype later — MannyKayy · 2026-09-11
- Maxinsights: The Data Supplier Behind Silicon Valley's Hottest Robots, With 2M+ Hours of Egocentric Data — 新智元 · 2026-09-11
- Inside China Opto-electronics Expo: The XR Supply Chain, From Waveguides to MicroLED — Scobleizer · 2026-09-11
- Tesla shows Cybercab in Japan, its steering wheel-free robotaxi built from scratch — yunta_tsai · 2026-09-11