Sony AI Open-Sources MIMO Iterative Audio Separation Framework That Works on Any Model

petewoodbridge · x · 2026-09-10

Sony AI released an architecture-agnostic framework that turns any single-step audio source separation model into an iterative refinement model while strictly preserving mixture consistency. By extending backbones to a multi-input multi-output (MIMO) configuration, all estimated stems are jointly refined per iteration. Applied to SCNet and BS-RoFormer, it yields cleaner karaoke vocal removal, better instrument isolation, and fewer leaks/artifacts. Paper and official PyTorch code are open-sourced — useful for producers, DJs, restoration, and music generation.

Related event: Sony AI open-sources MIMO iterative audio separation framework(2 posts)→

Original post →

More from Multimodal

Multimodal channel →