Alibaba Releases Real-Time Multimodal Model Wan-Streamer
anselm · x · 2026-07-11
Alibaba introduces Wan-Streamer, an end-to-end omni-modal understanding and generation model designed for real-time full-duplex interaction.
Key Capabilities
- A single model that can listen, see, understand, and respond while outputting synchronized audio and video.
- Built on a single native streaming Transformer, achieving a full video conversation latency of around 550ms.
- Unconstrained by fixed avatar skeletons, allowing it to instantly morph into various entities like humans, pets, or anime characters.
v0.2 Updates
- Resolution increased to 640×368.
- Frame rate set at 25 FPS.
- Model-side latency reduced to approximately 200ms.
More from Research
- Animation shows how an MLP’s first-layer weights change while learning MNIST — CatAstro_Piyush · 2026-07-22
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Project APE builds its verifier benchmark from 100 AI-written papers with injected errors — soumitrashukla9 · 2026-07-22
- Paper proposes a CRED taxonomy and benchmark to measure research-error detectors — soumitrashukla9 · 2026-07-22