Wan-Streamer: Video as World Plus Event Stream
Wan-AI · hf · 2026-07-17
Wan-Streamer v0.3: Treating Video as "World + Event Stream"
Wan-AI released Wan-Streamer v0.3, restructuring its real-time streaming interaction model with a unified perspective: Video = Stable World + Time-Varying Event Stream. The "world" includes relatively stable contexts like environments, scenes, characters, and audio; the "event stream" covers dynamic info like scene changes, character actions, speech, and other sounds.
Training Objective Brought by This Perspective
The model is trained to: given a world and continuous input, predict how this world changes, responds, and advances in real-time. The resulting capabilities can transfer to various real-time downstream tasks.
Specific Implementation
The authors applied it to real-time full-duplex audio-video interaction:
- Input: multimodal user content
- Output: speech in language form and behavioral actions
- The overall understanding process resembles a real-time version of vision-language-action
Performance/System Metrics
- Resolution: 640x368
- Frame rate: 25 FPS
- Streaming unit: 160 ms
- Model-side response latency: approx. 200 ms
- Under a 350 ms bidirectional network budget, total interaction latency is approx. 550 ms
This indicates it's not just a pure concept demo, but emphasizes a real-time interaction stack.
More from Multimodal
- A simple workflow to turn a photo into an image prompt using Gemini, Grok, or GPT Image — harshitagu72595 · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22
- Hand-painted figurines run through Seedance look eerily alive — cocktailpeanut · 2026-07-22
- An AI agent-made bayou country music video is making the rounds on Reddit — LazyKaleidoscope4696 · 2026-07-22
- Testing Qwen 3 Image: Map Borders Shift Based on Prompts, Includes Chinese Labels — NirantK · 2026-07-22
- Solo founder turns complaints on screen into bug reports with a local MCP server — phdptsd · 2026-07-22