Google and Tel Aviv researchers unveil SepGen, generating video with per-source stems for 4D spatial audio

YonatanBitton · x · 2026-10-11

Researchers from Tel Aviv University and Google present SepGen, which extends a pretrained audio-video generation model to emit one waveform per described source in a single pass, alongside the video and mix. It supports both generation (from text captions) and separation (splitting an existing video's audio into stems).

Because sources are individually accessible, they can be placed into a lifted 4D scene and rendered as spatial audio from novel viewpoints; the paper ships an interactive walk-through demo. Added weights leave the original soundtrack unchanged, a short extra training round boosts generation without hurting separation, and results are strongest on speech.

Original post →

More from Multimodal

Multimodal channel →