Xiaomi releases ControlFoley: video-to-audio generation model
apolinariosteps · x · 2026-08-19
Xiaomi released ControlFoley on Hugging Face, a video-to-audio model that generates sound based on video content. It features optional prompt input and audio reference guidance. A demo is available on Hugging Face Spaces.
More from Multimodal
- dots3-note preview launches with support for multi-modal long-horizon task planning — LearnWithBishal · 2026-08-19
- MiniMax H3 ecosystem roundup: Fizgig 4.1.2 trains character+voice LoRAs in one run — optimisticalish · 2026-08-19
- Detailed prompt structure for generating the Fabio goose meme video — danishkirel · 2026-08-19
- Cute animation transitions generated with MiniMax H3 — gizakdag · 2026-08-19
- Improving realistic fighting scenes with Minimax H3 and R2V — FreddyShrimp · 2026-08-19
- Fixing video color and details using first frame denoising — DuHal9000 · 2026-08-19