MiniMax Launches Music 3: Generates Up to 5-Minute Complete Songs
huggingface · x · 2026-08-14
MiniMax has released MiniMax Music 3, a high-performance music generation model, on Hugging Face. The model generates complete songs up to five minutes long based on lyrics and detailed music descriptions, maintaining long-range coherence, expressive vocals, and evolving arrangements.
Technical Architecture:
- Combines an 8B Global LLM for long-range musical structure and a 0.6B Local LLM for frame-level acoustic details.
- Uses a continuous hidden-state synthesis system based on Flow Matching and Flow-VAE.
- Outputs 32 kHz, 16-bit stereo WAV audio.
Fine-Grained Control: Accepts lyrics (with section tags like [Verse] and [Chorus]) and structured music descriptions (covering genre, BPM, key, vocal gender/timbre, etc.) for precise control.
Related event: MiniMax Open-Sources Music 3 Text-to-Music Model(14 posts)→
More from Models
- Vercel Offers GLM 5.2 Model Free for eve Agents Until August 27 — cramforce · 2026-08-14
- Deepgram Crosses $100M ARR and Launches Flux TTS Voice Model — deepgramscott · 2026-08-14
- SemiAnalysis: DeepMind Overhaul Signals Gemini's Downfall, GCP Emerges as Winner — ben_j_todd · 2026-08-14
- Musk Offers More Free Usage and Resets Limits for Grok 4.6 Launch — EricBuess · 2026-08-14
- Grok 4.6 Tops GPQA Diamond Leaderboard with 94.9% Score — elonmusk · 2026-08-14
- a16z's Martin Casado Tests Grok 4.6: Impressed by Complex Coding and Long Tasks — elonmusk · 2026-08-14