HelixWorld: Real-Time Audio-Visual World Model Runs at 24 FPS on a Single GPU
NoizAI · hf · 2026-10-05
HelixWorld is a real-time interactive audio-visual world model where visual scenes and camera-grounded spatial stereo sound co-evolve under user interaction, addressing the silence of prevailing world models.
- Curates a high-fidelity spatial audio-visual dataset with true stereo acoustics and metric camera poses; pre-trains a bidirectional teacher conditioned on 6-DoF camera trajectories and user actions
- Distills the teacher into a few-step streaming student via online trajectory distillation, sustaining drift-free joint audio-visual rollouts at 24 FPS on a single GPU
- Formalizes spatial-acoustic consistency and introduces HelixBench to evaluate whether synthesized sound fields track dynamic viewpoint motion
- Matches state-of-the-art silent world models in visual fidelity while significantly surpassing baselines in spatial-acoustic immersion
More from Multimodal
- Recreating the GPT semi-realistic 3D anime look locally with LoRAs, workflow and prompts — AI-Make-NSFW-Stuff · 2026-10-05
- WIP: using Claude to reconstruct structured Blender models from photos, beyond 3DGS — jwt0625 · 2026-10-05
- A reusable cinematic prompt: 1:5-scale miniature woman adventures in a convenience store — SimplyAnnisa · 2026-10-05
- UniEvo-VL: Multimodal Models Improve Image Generation via Self-Feedback — StanfordAILab · 2026-10-05
- BabyCast returns: testing Kling 4.0 Flash's talking-avatar skills in a 20-second challenge — CurieuxExplorer · 2026-10-05
- GPT-6.1 Sol + Blender MCP turns Gundam photos into a textured 3D model in under 10 minutes — sidahuj · 2026-10-05