Qwen Releases Technical Report for Qwen-Audio-3.0-Gen-Preview
udmrzn · x · 2026-07-31
The Alibaba Qwen team released the technical report for Qwen-Audio-3.0-Gen-Preview. The model uses a unified non-autoregressive framework combining a Diffusion Transformer (DiT) and a shared VAE to directly generate complete mixed waveforms containing speech, music, sound effects, and multiple roles.
A shared continuous VAE compresses 48kHz stereo waveforms into 25Hz latent sequences with semantic supervision. Evaluations show it achieves excellent speaker similarity on Seed-TTS-Eval, higher cross-turn consistency than Seed-Audio-1.0 on multi-speaker benchmarks, and stronger temporal localization on AudioCaps.
More from Multimodal
- Creator showcases short film generated with Runway Seedance 2.0 — Lucidjordan79 · 2026-07-31
- Can One LoRA Hold Multiple Concepts? Devs Discuss Multi-Style Training Techniques — Lounlysoul007 · 2026-07-31
- Fish Audio Raises $52M Seed, Launches S2.1 Pro Voice Model — thisdudelikesAI · 2026-07-31
- KalpaLabs Launches Conversational Speech Model in Public Beta — ycombinator · 2026-07-31
- FLUX 3 Preview Goes Live; Nous Research Launches Short Film Contest — NousResearch · 2026-07-31
- Beginner Question: How to Handle Dataset Captions for LoRA Training? — Mean-Crab1827 · 2026-07-31