Google and Tel Aviv researchers unveil SepGen, generating video with per-source stems for 4D spatial audio
YonatanBitton · x · 2026-10-11
Researchers from Tel Aviv University and Google present SepGen, which extends a pretrained audio-video generation model to emit one waveform per described source in a single pass, alongside the video and mix. It supports both generation (from text captions) and separation (splitting an existing video's audio into stems).
Because sources are individually accessible, they can be placed into a lifted 4D scene and rendered as spatial audio from novel viewpoints; the paper ships an interactive walk-through demo. Added weights leave the original soundtrack unchanged, a short extra training round boosts generation without hurting separation, and results are strongest on speech.
More from Multimodal
- Testing spatial-temporal consistency: recreating one moment from two views on Kling 4.0 Flash — umesh_ai · 2026-10-11
- Open-source art animation skill offers 35 art styles and 9 narration grammars — AlchainHust · 2026-10-11
- Qwen-Image Edit lands native support in ComfyUI — saroxel · 2026-10-11
- AI video made with Seedance 2.5 hailed as a masterpiece — SimplyAnnisa · 2026-10-11
- Anthropic launches Claude Motion beta: turn reports into editable MP4 animated explainers — PrajwalTomar_ · 2026-10-11
- Sunday Video Challenge #85: Create a Surreal 'Metamorphosis' AI Video in 30 Seconds — LudovicCreator · 2026-10-11