Making a music video with Minimax H3: a full-reference video prompt workflow
Select_Bowler3099 · reddit · 2026-09-08
A Redditor shares the complete workflow behind their first music video made with Minimax H3: R2V + a 4-step LoRA at 0.6 MP, upscaled via SeedVR2.
The centerpiece is a "video prompt engineer" system prompt for full-reference models:
- Fixed six-section output: subjectdefinitions, summary, retentionanalysis, detaileddescription, overallsoundscape, nondiegeticmusic.
- Consistent <Subject N> / <Picture N> / <Video N> / <Audio N> labels referencing people, keyframes, videos, and audio tracks across all sections.
- summary must start with a task-type prefix like [reference generation] or [audio reference].
- Strict faithfulness rules: no added/omitted visual elements or actions, dialogue kept in <d> tags in the original language.
The template is transferable to any video model supporting full-reference mode.
More from Multimodal
- ByteDance secretly building real-time world model on Seedance, October launch rumored — dotey · 2026-09-09
- Redditor vibe-codes free CMYK Ben Day dots node for ComfyUI with 37 presets — darkside1977 · 2026-09-09
- AI filmmaking shifts from describing shots to defining worlds, with a GPT-6 Astra to video workflow — AIwithGhotai · 2026-09-09
- 'Regenerate' is becoming the Ctrl+Z of AI video: fix the broken part instead — Fast_Dependent742 · 2026-09-09
- Inworld's TTS model hits GA as CEO explains its end-to-end voice stack — agihouse_org · 2026-09-09
- GPT-6 Astra codes geometry, Blender + Clay plugin renders, Dreamina finalizes footage — thetripathi58 · 2026-09-09