Qwen-Music Music Generation Model Report

Qwen · hf · 2026-07-20

Alibaba released a technical report for Qwen-Music, a music generation model capable of producing high-quality songs with full vocal performances. It supports two types of tasks: text-to-music generation, and covering existing songs with different styles or timbres.

The model architecture consists of three parts: Qwen-Music-Tokenizer compresses audio into 25Hz single-codebook music semantic tokens; Qwen-Music-LLM performs autoregressive generation based on these tokens, introducing Melody-CoT to plan the melody before generating the entire song; and Qwen-Music-Render handles generative stereo rendering to fill in audio quality details. For training, the team used over 5 million hours of multilingual music data, combined with quality-aware pre-training, supervised initialization, offline DPO, and online GSPO. The report claims Qwen-Music achieves SOTA in 13 out of 16 objective music/audio metrics and is preferred by professional reviewers. Compared to mainstream closed-source systems, it also better preserves reference melodies during covers.

Original post →

More from Models

Models channel →