MiniMax Music 3 Surfaces on Hugging Face, Supports 5-Minute Audio
NerdyRodent · x · 2026-08-14
MiniMax has released the MiniMax Music 3 model on Hugging Face.
- Core capabilities: Natively supports full-song generation up to five minutes, maintaining musical themes, rhythm, vocal identity, and arrangement progression across long sequences.
- Architecture: Combines an 8B Global LLM for long-range musical structure and a 0.6B Local LLM for frame-level acoustic detail, using a continuous hidden-state synthesis system based on Flow Matching and Flow-VAE.
- Fine-grained control: Accepts lyrics with section tags and structured music descriptions (genre, BPM, timbre, emotional progression) for precise control, outputting 32 kHz, 16-bit stereo WAV audio.
Related event: MiniMax Releases Open-Source Music 3 Music Generation Model(17 posts)→
More from Multimodal
- Help Needed: How to build a stable MiniMax R2V workflow in ComfyUI? — haremlifegame · 2026-08-14
- Opus 5 and Thrixel One-Shot a Playable Browser Flight Game — RanaHanocka · 2026-08-14
- CapCut Launches Seedance 2.5 Globally with 1080p Video Continuation Challenge — Aiden_Tech_Ai · 2026-08-14
- Automating ComfyUI Bulk Generation with Python and Free Gemini API — Excellent_Scene7402 · 2026-08-14
- Flux 3 Text-to-Video Probe Shows Impressive Realism and Camera Control — gen_ericai · 2026-08-14
- H3 Model Tested: Generates 10s 480p Video in 10 Mins on 12GB VRAM — cocktailpeanut · 2026-08-14