HOMIE improves human-object video personalization with multimodal alignment
Yiyang Cai · hf · 2026-07-21
HOMIE is a multimodal framework for human-object centric video personalization.
- The paper targets subject-driven video generation, where existing methods struggle to balance subject fidelity with correct human-object interactions, especially when the object is an abstract concept such as a logo.
- It also addresses intra-subject references such as OCR maps or multi-view inputs, which are supposed to improve fidelity but are often not well aligned by prior methods.
- HOMIE uses a better MLLM integration strategy, adds global multimodal guidance inside self-attention to align semantic features with VAE tokens, and introduces modality-reference embeddings to separate MLLM features from VAE tokens and link reference-image tokens.
- The authors report state-of-the-art results across multiple HOCVP tasks.
Related event: HOMIE Enables Human-Object Video Personalization(2 posts)→
More from Multimodal
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Non-coder builds full-featured Android ComfyUI client with ChatGPT, submits to Google Play — ComfierUI · 2026-09-11
- FastH3-Live hits 22fps: acceleration node benchmarks and the --vram-headroom trick — spartong945 · 2026-09-11
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11