Alibaba Releases Wan3.0 Video Model with Native 30-Second Generation
Alibaba has officially released the Wan3.0 video generation model and opened it for public testing. The new version achieves a breakthrough in generating up to 30 seconds of video in a single pass, featuring comprehensive upgrades in generation duration, multimodal inputs, and visual realism. This model sets a new standard in the video generation space and warrants close attention.
Confirmed
- Core Capabilities: Wan3.0 natively supports generating videos up to 30 seconds long, featuring reality-level rendering with excellent performance in understanding physics, lighting details, and character realism.
- Cinematography & Narrative: Supports continuous single-take shots and director-level camera movements (such as montage), ensuring more coherent and logical storytelling.
- Universal Reference Inputs: Moving beyond traditional text, image, audio, and video inputs, it now supports multimodal formats like documents, spreadsheets, and slides.
Why It Matters
- With generation duration breaking through to a native 30 seconds, the narrative potential for a single generation is vastly expanded. Coupled with multimodal reference inputs, this further broadens the possibilities for applying AI video in real-world scenarios.
2026-08-06 ~ 2026-08-06 · 5 related posts
Primary sources
- Alibaba's Wan3.0 Launches Public Beta: Native 30-Second Video Generation — bdsqlsz · 2026-08-06
- [source] Alibaba Launches Wan3.0 Video Model: Generates 30-Second Continuous Shots — 千问APP · 2026-08-06
- Alibaba Releases Wan3.0: Native 30s Video and Multi-Format Inputs — xiaohu · 2026-08-06
- Alibaba Launches WAN 3.0: 30-Second Video Generation with Multimodal References — aziz4ai · 2026-08-06
1 near-duplicate retellings: xiaohu