HOMIE improves human-object video personalization with multimodal alignment
Yiyang Cai · hf · 2026-07-21
HOMIE is a multimodal framework for human-object centric video personalization. - The paper targets subject-driven video generation, where existing methods struggle to balance subject fidelity with correct human-object interactions, especially when the object is an abstract concept such as a logo. - It also addresses intra-subject references such as OCR maps or multi-view inputs, which are supposed to improve fidelity but are often not well aligned by prior methods. - HOMIE uses a better MLLM integration strategy, adds global multimodal guidance inside self-attention to align semantic features with VAE tokens, and introduces modality-reference embeddings to separate MLLM features from VAE tokens and link reference-image tokens. - The authors report state-of-the-art results across multiple HOCVP tasks.
More from Multimodal
- Seedance 2.0 turns one reference image into a cinematic fight scene — techhalla · 2026-07-21
- Seedance 2.0 keeps character consistency across 15+ shots with just 3 prompts — techhalla · 2026-07-21
- DecartAI’s Lucy 2.5 Realtime lands on fal with live video-to-video editing — gorkem · 2026-07-21
- Google Gemini’s Omni text-to-video output is getting better, user says — michaelrabone · 2026-07-21
- OpenArt AI demos a Video Remix tool that can transform an existing video — eyishazyer · 2026-07-21
- ElevenLabs raises ElevenMusic free usage to 400 tracks a month — lukeharries · 2026-07-21