Testing MiniMax H3: Generating 30s coherent music with structured prompts

-Ellary- · reddit · 2026-08-07

A developer shared a hands-on test of using the MiniMax H3 model as a music generation engine. Using structured prompts, H3 can generate up to 30 seconds of coherent audio, supporting custom lyrics, genres, and instrument arrangements.

The author provided a prompt example for a 1990s hip-hop rap style, detailing the timeline (e.g., [0s] intro, [5s] vocal entry, [25s] heavy outro drop) to demonstrate precise control over the song's structure and emotional build-up.

Original post →

More from Multimodal

Multimodal channel →