HelixWorld 1.0: First real-time interactive audio-video world model
量子位 · wechat · 2026-08-17
NoizAI, in collaboration with researchers from HKUST, Tsinghua, CMU, and Google DeepMind, has released HelixWorld 1.0, the first real-time interactive audio-video world model. It generates continuous visuals and 48kHz binaural audio at 24FPS based on user input and camera movements. Unlike traditional dubbing, it uses a unified Transformer to generate audio and video simultaneously, ensuring spatial consistency. The team utilized a million-scale aligned dataset, action-conditioned joint generation, and causal inference with KVCache for real-time performance. Weights and code will be fully open-sourced.
More from Multimodal
- Frameo tops Physion Labs' AI filmmaking agent benchmark by 20 points — NirantK · 2026-08-17
- FARÖ: an AI-made mini-series where heartbreak and reunion share identical midsummer light — gen_ericai · 2026-08-17
- MiniMax H3 on Magnific supports 2K video with native synced sound — aftahi_ai · 2026-08-17
- Scanner aesthetic: Creating flatbed-scan portraits with Midjourney — tisch_eins · 2026-08-17
- Looking for consistent, unfiltered AI image generator with reference upload — Key_Poetry_7829 · 2026-08-17
- Challenges with coherence and camera control using Minimax H3 — BoneDaddyMan · 2026-08-17