Sony AI Open-Sources MIMO Iterative Audio Separation Framework That Works on Any Model
petewoodbridge · x · 2026-09-10
Sony AI released an architecture-agnostic framework that turns any single-step audio source separation model into an iterative refinement model while strictly preserving mixture consistency. By extending backbones to a multi-input multi-output (MIMO) configuration, all estimated stems are jointly refined per iteration. Applied to SCNet and BS-RoFormer, it yields cleaner karaoke vocal removal, better instrument isolation, and fewer leaks/artifacts. Paper and official PyTorch code are open-sourced — useful for producers, DJs, restoration, and music generation.
Related event: Sony AI open-sources MIMO iterative audio separation framework(2 posts)→
More from Multimodal
- MiniMax H3 video demo shows smooth live-action footage with dynamic typography — Hailuo_AI · 2026-09-11
- GPT image outputs upscaled 10x to 100MP via free tool, cookie trick extends quota — vista8 · 2026-09-10
- Reddit user shares first narrative AI short film "To Nora" — Individual_Item_3140 · 2026-09-10
- The Prop Sheet Everyone Skips: One Generation Keeps AI Film Endings From Hallucinating — socialwithaayan · 2026-09-10
- Studios Spend Millions on AI Animation; This Creator Says One Afternoon Suffices — socialwithaayan · 2026-09-10
- A simple Windows WebUI for the open-source music model YuE2 is now on GitHub — LadyQuacklin · 2026-09-10