ECCV Paper: Step-by-Step Video-to-Audio Synthesis via Negative Guidance

mittu1204 · x · 2026-08-31

Inspired by traditional Foley workflows, this paper proposes a step-by-step video-to-audio (V2A) method allowing incremental sound event authoring. To avoid costly multi-reference datasets, each step uses negative guidance to suppress sounds from previous tracks. The guidance model is fine-tuned on non-overlapping segments of standard single-reference datasets, leveraging acoustic context while staying visually grounded. Evaluations show improved sound separability and composite audio quality.

Original post →

More from Multimodal

Multimodal channel →