Peking Univ & Kling Unveil MAVIN to Solve Multi-Shot Video Narrative Challenges
机器之心 · wechat · 2026-08-28
Peking University, in collaboration with the Kling team and others, introduced MAVIN, a multi-shot audio-video generation model designed to address temporal misalignment and identity confusion in narrative video generation. MAVIN utilizes boundary-aware attention and identity-aware propagation mechanisms to synchronize cuts, lip movements, and character identities via structured scripts. The team also built the MAVINSet dataset with hierarchical annotations. Experiments show MAVIN achieves state-of-the-art results on 13 metrics and supports applications like narrative rhythm control and identity customization.
More from Multimodal
- Omni, Qwen, and other Flash models tested; video generated under 10s — iScienceLuvr · 2026-08-28
- Midjourney v8.2 Editing Features Tested: Removal & Repaint — LudovicCreator · 2026-08-28
- Krea AI launches new Krea Three model at NYC event — chrisfirst · 2026-08-28
- Google DeepMind Releases ORBIT++ Benchmark for SfM in 360° Video — taiyasaki · 2026-08-28
- MiniMax H3 pure text-to-video test with a full shot-by-shot prompt — BitterAd8431 · 2026-08-28
- Midjourney tests V8.2 edit model with inpainting, outpainting, and image-to-image — DavidSHolz · 2026-08-28