WorldSonus adds real-time spatial stereo audio to world models at 0.41 real-time factor
NoizAI · hf · 2026-10-08
WorldSonus is an interactive video-to-audio framework for world models, tackling real-time generation, mid-stream interactive sound control, and spatially aligned stereo. It uses a streaming causal autoregressive diffusion architecture with a low real-time factor of 0.41, an audio-centric captioning pipeline with chunk-indexed prompt scheduling for dynamic sound manipulation, and stereo supervision from ambisonic data for spatial alignment. It matches or outperforms state-of-the-art bidirectional models on open-domain video-to-audio benchmarks.
More from Multimodal
- AI-generated feature 'A Woman Asleep' enters major film festival's main competition — lmoroney · 2026-10-08
- vLLM-Omni technical report: a unified serving runtime for omni-modal generation — vllm_project · 2026-10-08
- Alaskan Raven Couple 'Conversing' Video Goes Viral on X — ZeroStateReflex · 2026-10-08
- Band Builds Audio-Reactive WebGL + Local SD 1.5 Pipeline for Live Improv Music Video — XploitXploit · 2026-10-08
- Claude turns OpenAI's 198-page Erdős conjecture proof into a 2-minute narrated 3D animation — imjustnewatai · 2026-10-08
- AI Short Film About Relationships Made With ComfyUI Agent Driver and Multi-Model Pipeline — TheHollywoodGeek · 2026-10-08