Qwen-Music Music Generation Model Report
Qwen · hf · 2026-07-20
Alibaba released a technical report for Qwen-Music, a music generation model capable of producing high-quality songs with full vocal performances. It supports two types of tasks: text-to-music generation, and covering existing songs with different styles or timbres.
The model architecture consists of three parts: Qwen-Music-Tokenizer compresses audio into 25Hz single-codebook music semantic tokens; Qwen-Music-LLM performs autoregressive generation based on these tokens, introducing Melody-CoT to plan the melody before generating the entire song; and Qwen-Music-Render handles generative stereo rendering to fill in audio quality details. For training, the team used over 5 million hours of multilingual music data, combined with quality-aware pre-training, supervised initialization, offline DPO, and online GSPO. The report claims Qwen-Music achieves SOTA in 13 out of 16 objective music/audio metrics and is preferred by professional reviewers. Compared to mainstream closed-source systems, it also better preserves reference melodies during covers.
More from Models
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11