UMO: unified in-context framework turns text2motion model into multi-task motion foundation

Michael_J_Black · x · 2026-09-11

UMO (Brown, MIT, Max Planck, Meta Reality Lab, HKU; ECCV 2026) builds on a pre-trained 3D text-to-motion model and uses a unified in-context learning framework with three meta-operations to handle diverse motion tasks — in-domain text-to-motion generation plus out-of-domain temporal tasks like keyframe infilling, prediction, backcasting and in-betweening — beating task-specific baselines. Paper and code are public.

Original post →

More from Embodied

Embodied channel →