Peking Univ & Kling Unveil MAVIN to Solve Multi-Shot Video Narrative Challenges

机器之心 · wechat · 2026-08-28

Peking University, in collaboration with the Kling team and others, introduced MAVIN, a multi-shot audio-video generation model designed to address temporal misalignment and identity confusion in narrative video generation. MAVIN utilizes boundary-aware attention and identity-aware propagation mechanisms to synchronize cuts, lip movements, and character identities via structured scripts. The team also built the MAVINSet dataset with hierarchical annotations. Experiments show MAVIN achieves state-of-the-art results on 13 metrics and supports applications like narrative rhythm control and identity customization.

Original post →

More from Multimodal

Multimodal channel →