MAVIN: multi-shot audio-video generation with narrative control, ECCV 2026 Oral

jiqizhixin · x · 2026-09-06

Peking University, Kling, CASIA, and Sun Yat-sen University present MAVIN, a multi-shot audio-visual generation model with customized narrative control, accepted as an ECCV 2026 Oral.

It tackles three core challenges:

The model respects future-shot descriptions without semantic leakage. Shot transition accuracy (STA) reaches 0.98.

Original post →

More from Multimodal

Multimodal channel →