Making an AI Music Video With MiniMax H3: Upgrading the Model Actually Made It Worse
WolframRvnwlf · x · 2026-09-30
A creator shares lessons from producing a fully AI-generated music video:
- Bigger model ≠ better video: Moving from MiniMax H3 Fused 4-Step to Full didn't automatically help — some shots gained detail but lost the camera movement or performance they liked, plus extra background people, musicians singing when they shouldn't, and broken instruments. They rejected the full replacement and picked the best shots across old and new versions.
- Tiny audio details caused the longest fights with the machines: Whisper large-v3 and Gemini 3.5 Flash disagreed on sung words and the number of repeated "WE BITE BACK" calls. Transcripts were useful clues but unreliable for precise lyric timing — that needed listening, waveform inspection and musical context.
- A separated, time-adjusted vocal repair sounded wrong; the better fix reused an intact phrase from an earlier chorus, aligned.
Stack: ACE-Step 1.5 XL SFT LM 4B (song/instruments/vocals), MiniMax H3 Ref2VA (reference-guided performances), FLUX.2 [klein] 9B (musician references), ByteDance's Seedream 5.0 Pro and SeedVR2 (stage reference and full-band shots).
More from Multimodal
- NVIDIA's LongLive-Plug: Distill Once, Deploy Training-Free Across 54 Downstream Video Models — nvidia · 2026-09-30
- Adobe Research Shows Adversarial Post-Training Restores Missing High-Frequency Detail in Pixel Diffusion — adobe-research · 2026-09-30
- UCSD's LIFT Lets You Control Future Video Layouts via On-Policy Self-Distillation — UCSanDiego · 2026-09-30
- MiniMax-H3 RefMod Upgrade: Packing JPEGs Into Safetensors Cuts Encoding Time and Kills Identity Bleed — acedelgado · 2026-09-30
- Creator Says Opus 5.5 Renders in 45 Mins What Used to Need Four-Figure Plugins — justin_hart · 2026-09-30
- Casey: AI video will inevitably get so good that critics must admit great works — johncoogan · 2026-09-30