ECCV 2026: Odoriko Enables Shape-Aware Unified Multimodal Motion Generation

mittu1204 · x · 2026-08-31

Existing unified multimodal motion frameworks ignore morphological factors like gender and body shape, treating all subjects as equivalent. Odoriko is the first unified framework (text, music, video) to incorporate bio-morphological information, generating motion consistent with "who" is moving. When explicit morphology is unavailable, it jointly recovers shape and motion. Experiments show it matches or exceeds specialized models across benchmarks while enabling morphology-consistent generation.

Original post →

More from Multimodal

Multimodal channel →