First systematic survey of joint video-audio generation and editing: 9 categories, 28 edit types
Abhinav Sharma · hf · 2026-10-02
A new survey systematically examines generative methods that model video and audio jointly, noting that although the two modalities are perceived together, most generative models treat them in isolation.
Highlights:
- A unified formulation casts joint generation, cross-modal generation, and joint editing as three problems on a single distribution over audio-visual pairs.
- A design taxonomy compares methods along five axes centered on one question: how is output kept coherent across modalities in time and semantics?
- Claimed as the first overview to taxonomize joint audio-visual editing: nine edit categories spanning 28 edit types.
- Covers methods, datasets, and metrics per setting, closing with the most consequential open problems.
A useful roadmap for researchers in multimodal generation.
More from Multimodal
- HuST Lab's Multimodal Flow: Fully Continuous Unified Language-Vision Generation — hustvl · 2026-10-02
- Tencent Hunyuan's Tex-Zero Trains 3D Texture Generation Without Any 3D Assets — Tencent-Hunyuan · 2026-10-02
- One prompt gets Claude Fable 5.5 to render a AAA-grade 3D Rube Goldberg machine — imjustnewatai · 2026-10-02
- Seedance 2.5 Video Goes Viral for Ultra-Realistic Early-2000s DV Camcorder Look — SimplyAnnisa · 2026-10-02
- ComfyUI Clips + sync.so Lip-Sync: Edit Before or After Processing? — Positive_Society1876 · 2026-10-02
- Dolly or Zoom? Prompting AI Video for Camera-Only Motion on a Static Car — Classic-Jellyfish-26 · 2026-10-02