HelloWorld: Enabling Real-Time Social Interaction in Video World Models
Liangyang Ouyang · hf · 2026-08-06
Current video world models lack social interaction between users and in-world characters. HelloWorld introduces a model allowing users to trigger on-screen character responses (e.g., waving, speaking) via a button press.
- Self-Distillation: Finetunes the model on self-synthesized data containing both social interactions and camera motion without degrading quality.
- Training-Free Inference: Modulates DiT cross-attention masks upon button press to temporally localize character responses.
- Benchmark: Introduces HelloWorldBench (400 samples). Experiments show state-of-the-art picture aesthetics alongside superior interaction quality.
More from Multimodal
- Advanced MiniMax H3 Filmmaking Workflow: From Prompts to Editing — Tricky_Algae2625 · 2026-08-06
- Testing LTX upscaler: Generating high-res images on low VRAM GPUs — AniZeee · 2026-08-06
- Elon Musk Announces Launch of Grok Imagine Image Generation — elonmusk · 2026-08-06
- Meta Unpacks Multimodal Pretraining: Strong Generation with 5% Compute — facebook · 2026-08-06
- UniWorld-View: Large-Baseline Novel View Synthesis via Video Diffusion — Haiyang Zhou · 2026-08-06
- ToolArtist: Agentic Image Generation via Unified Multimodal Models — Jiahao Zhao · 2026-08-06