HOMIE improves human-object video personalization with multimodal alignment

Yiyang Cai · hf · 2026-07-21

HOMIE is a multimodal framework for human-object centric video personalization. - The paper targets subject-driven video generation, where existing methods struggle to balance subject fidelity with correct human-object interactions, especially when the object is an abstract concept such as a logo. - It also addresses intra-subject references such as OCR maps or multi-view inputs, which are supposed to improve fidelity but are often not well aligned by prior methods. - HOMIE uses a better MLLM integration strategy, adds global multimodal guidance inside self-attention to align semantic features with VAE tokens, and introduces modality-reference embeddings to separate MLLM features from VAE tokens and link reference-image tokens. - The authors report state-of-the-art results across multiple HOCVP tasks.

Original post →

More from Multimodal

Multimodal channel →