Qwen Releases Technical Report for Qwen-Audio-3.0-Gen-Preview

udmrzn · x · 2026-07-31

The Alibaba Qwen team released the technical report for Qwen-Audio-3.0-Gen-Preview. The model uses a unified non-autoregressive framework combining a Diffusion Transformer (DiT) and a shared VAE to directly generate complete mixed waveforms containing speech, music, sound effects, and multiple roles.

A shared continuous VAE compresses 48kHz stereo waveforms into 25Hz latent sequences with semantic supervision. Evaluations show it achieves excellent speaker similarity on Seed-TTS-Eval, higher cross-turn consistency than Seed-Audio-1.0 on multi-speaker benchmarks, and stronger temporal localization on AudioCaps.

Original post →

More from Multimodal

Multimodal channel →