WorldSonus adds real-time spatial stereo audio to world models at 0.41 real-time factor

NoizAI · hf · 2026-10-08

WorldSonus is an interactive video-to-audio framework for world models, tackling real-time generation, mid-stream interactive sound control, and spatially aligned stereo. It uses a streaming causal autoregressive diffusion architecture with a low real-time factor of 0.41, an audio-centric captioning pipeline with chunk-indexed prompt scheduling for dynamic sound manipulation, and stereo supervision from ambisonic data for spatial alignment. It matches or outperforms state-of-the-art bidirectional models on open-domain video-to-audio benchmarks.

Original post →

More from Multimodal

Multimodal channel →