MAVIN: multi-shot audio-video generation with narrative control, ECCV 2026 Oral
jiqizhixin · x · 2026-09-06
Peking University, Kling, CASIA, and Sun Yat-sen University present MAVIN, a multi-shot audio-visual generation model with customized narrative control, accepted as an ECCV 2026 Oral.
It tackles three core challenges:
- Temporal misalignment (who speaks when, no voice bleeding across cuts)
- Identity confusion: per-character visual and voice anchors keep faces and voices consistent across shots in multi-person dialogue
- Incomplete input: turning a synopsis into an executable shot script
The model respects future-shot descriptions without semantic leakage. Shot transition accuracy (STA) reaches 0.98.
More from Multimodal
- LLaDA edit turbo: new image editing model from the diffusion-LLM project — Slight_Tone_2188 · 2026-09-06
- Testing the 'DLSS 5' image workflow: 29s per image on a 5060 Ti — thatguyjames_uk · 2026-09-06
- Midjourney tip: strip facial detail to test if silhouette carries the emotion — tisch_eins · 2026-09-06
- Two-person team spent 12 days and $400+ credits making a 6-min AI short film with Codex + Seedance 2.5 — APPSO · 2026-09-06
- Surreal AI video hides a tiny train station under a footprint, full prompt shared — umesh_ai · 2026-09-06
- uncomfymcp: a minimal ComfyUI MCP server for text-to-image chat workflows without exporting APIs — citrainmyhefeweizen · 2026-09-06