Wan-Streamer: Video as World Plus Event Stream
Wan-AI · hf · 2026-07-17
Wan-Streamer v0.3: Treating Video as "World + Event Stream"
Wan-AI released Wan-Streamer v0.3, restructuring its real-time streaming interaction model with a unified perspective: Video = Stable World + Time-Varying Event Stream. The "world" includes relatively stable contexts like environments, scenes, characters, and audio; the "event stream" covers dynamic info like scene changes, character actions, speech, and other sounds.
Training Objective Brought by This Perspective
The model is trained to: given a world and continuous input, predict how this world changes, responds, and advances in real-time. The resulting capabilities can transfer to various real-time downstream tasks.
Specific Implementation
The authors applied it to real-time full-duplex audio-video interaction:
- Input: multimodal user content
- Output: speech in language form and behavioral actions
- The overall understanding process resembles a real-time version of vision-language-action
Performance/System Metrics
- Resolution: 640x368
- Frame rate: 25 FPS
- Streaming unit: 160 ms
- Model-side response latency: approx. 200 ms
- Under a 350 ms bidirectional network budget, total interaction latency is approx. 550 ms
This indicates it's not just a pure concept demo, but emphasizes a real-time interaction stack.
More from Multimodal
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- MiniMax Music Production Toolkit 2.5 for ComfyUI adds full mastering chain — Vivid_Promise1700 · 2026-09-11
- New Node Finder for ComfyUI ranks fresh nodes by star velocity and recency — Luke2642 · 2026-09-11
- Using a finisher move on one mosquito with MiniMax H3 MAX — the bug survives — Hailuo_AI · 2026-09-11