Qwen-Music Music Generation Model Report
Qwen · hf · 2026-07-20
Alibaba released a technical report for Qwen-Music, a music generation model capable of producing high-quality songs with full vocal performances. It supports two types of tasks: text-to-music generation, and covering existing songs with different styles or timbres.
The model architecture consists of three parts: Qwen-Music-Tokenizer compresses audio into 25Hz single-codebook music semantic tokens; Qwen-Music-LLM performs autoregressive generation based on these tokens, introducing Melody-CoT to plan the melody before generating the entire song; and Qwen-Music-Render handles generative stereo rendering to fill in audio quality details. For training, the team used over 5 million hours of multilingual music data, combined with quality-aware pre-training, supervised initialization, offline DPO, and online GSPO. The report claims Qwen-Music achieves SOTA in 13 out of 16 objective music/audio metrics and is preferred by professional reviewers. Compared to mainstream closed-source systems, it also better preserves reference melodies during covers.
More from Models
- Bug Hunt Bench ranks frontier coding models on 105 planted real-repo bugs — PawelHuryn · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- 105 hidden bugs, 2 repos: DeepSeek V4.1 Flash fixes 24 at $1.80 vs Opus 5's 27 at $51.33 — ChartsJournalX · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11