ECCV Paper: Step-by-Step Video-to-Audio Synthesis via Negative Guidance
mittu1204 · x · 2026-08-31
Inspired by traditional Foley workflows, this paper proposes a step-by-step video-to-audio (V2A) method allowing incremental sound event authoring. To avoid costly multi-reference datasets, each step uses negative guidance to suppress sounds from previous tracks. The guidance model is fine-tuned on non-overlapping segments of standard single-reference datasets, leveraging acoustic context while staying visually grounded. Evaluations show improved sound separability and composite audio quality.
More from Multimodal
- MiniMax H3 Outperforms Seedance 2.5 in Ad Gen Test — aziz4ai · 2026-08-31
- Jason & His Ai'Band release MV "The Noise of Our Lives" — Kyrannio · 2026-08-31
- LTX 2.5 tiled workflow enables 8K video upscaling — makeitrad1 · 2026-08-31
- MiniMax h3 tested for educational videos handles infographic motion shockingly well — mrcheriftas · 2026-08-31
- MiniMax H3 Prompt Writer adds Windows standalone build and Qwen GGUF support — nnorbbi · 2026-08-31
- MiniMax image tip: high megapixels + low steps beats low MP + many steps — Naruwashi · 2026-08-31