Qwen-Music Music Generation Model Report
Qwen · hf · 2026-07-20
Alibaba released a technical report for Qwen-Music, a music generation model capable of producing high-quality songs with full vocal performances. It supports two types of tasks: text-to-music generation, and covering existing songs with different styles or timbres.
The model architecture consists of three parts: Qwen-Music-Tokenizer compresses audio into 25Hz single-codebook music semantic tokens; Qwen-Music-LLM performs autoregressive generation based on these tokens, introducing Melody-CoT to plan the melody before generating the entire song; and Qwen-Music-Render handles generative stereo rendering to fill in audio quality details. For training, the team used over 5 million hours of multilingual music data, combined with quality-aware pre-training, supervised initialization, offline DPO, and online GSPO. The report claims Qwen-Music achieves SOTA in 13 out of 16 objective music/audio metrics and is preferred by professional reviewers. Compared to mainstream closed-source systems, it also better preserves reference melodies during covers.
More from Models
- A joking post asks whether this was the famous “move 37” moment for math — NielsRogge · 2026-07-21
- Xiaohongshu’s dots-note-3.0 gets a perfect IMO score and becomes the world’s second gold model — 量子位 · 2026-07-21
- Reddit user says Grok 4.5 felt faster and better than Claude for coding workflows — Rare_Iron9142 · 2026-07-21
- Kimi and GLM distillation debate reignites over what counts as real innovation — basedjensen · 2026-07-21
- Meme mocks Google’s AI lead after early Gemini 3.6 Flash outputs look rough — max_paperclips · 2026-07-21
- Users say GPT-5.6 Ultra feels like extra token burn with little visible gain — CtrlAltDwayne · 2026-07-21