Wan-Streamer: Treating Video as World + Event Stream
bronzeagepapi · x · 2026-07-18
Alibaba's Wan team has introduced Wan-Streamer v0.3, which conceptualizes video uniformly as a "World + Event Stream".
- World: Relatively stable context such as scenes, characters, audio, and environments.
- Event Stream: Time-varying content like speech, actions, gestures, and sound reactions.
With this definition, the model treats "predicting how the world moves, changes, and responds in real-time given the world and inputs" as a general pre-training task. The paper highlights its focus on real-time full-duplex audio-visual interaction, treating the event stream as the agent's language, behavior, and action output to achieve low-latency synchronized generation. Reported metrics include 640×368 resolution, 25 FPS, a 160ms streaming unit, 200ms model-side response latency, and 550ms total interaction latency (under a 350ms bidirectional network budget).
More from Multimodal
- Gemini Omni Flash turns a boat cabin into a cave in Flow by Google — chrisfirst · 2026-07-22
- A simple workflow to turn a photo into an image prompt using Gemini, Grok, or GPT Image — harshitagu72595 · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22
- Hand-painted figurines run through Seedance look eerily alive — cocktailpeanut · 2026-07-22
- An AI agent-made bayou country music video is making the rounds on Reddit — LazyKaleidoscope4696 · 2026-07-22
- Testing Qwen 3 Image: Map Borders Shift Based on Prompts, Includes Chinese Labels — NirantK · 2026-07-22