HelixWorld 1.0: First real-time interactive audio-video world model

量子位 · wechat · 2026-08-17

NoizAI, in collaboration with researchers from HKUST, Tsinghua, CMU, and Google DeepMind, has released HelixWorld 1.0, the first real-time interactive audio-video world model. It generates continuous visuals and 48kHz binaural audio at 24FPS based on user input and camera movements. Unlike traditional dubbing, it uses a unified Transformer to generate audio and video simultaneously, ensuring spatial consistency. The team utilized a million-scale aligned dataset, action-conditioned joint generation, and causal inference with KVCache for real-time performance. Weights and code will be fully open-sourced.

Original post →

More from Multimodal

Multimodal channel →