Qwen Releases UniSwap: Streaming Audio-Visual Identity Swapping Model

QwenBusinessUnit · hf · 2026-08-14

Alibaba's Qwen team released the UniSwap model, designed for talking videos to achieve synchronized appearance and voice replacement.

UniSwap employs a unified streaming audio-visual diffusion transformer architecture, coupled with specialized training and inference adaptations to maintain high-quality audio-visual sync and identity transformation during video generation.

Original post →

More from Multimodal

Multimodal channel →