One-Man Band Workflow: Producing an AI Music Video with Local Minimax
lazyspock · reddit · 2026-08-10
A creator detailed the full workflow of using the local Minimax model to produce a complete music video for an original song.
- Asset Preparation: Lyrics were self-written, while music/vocals were generated via Suno. The virtual singer was created using Z-Image and Krea2 for the base face, Flux for additional angles, and Ostris to train a LoRA.
- Video Generation: The song was sliced into 5-15 second clips, and prompts were generated with a custom GPT based on Minimax docs. Prompts were tested locally on a 4070 GPU at 0.3MP before final 1280x736 rendering on a rented Runpod 5090.
- Cost & Post-production: The project involved 21 shots and 30-35 renders, costing $12-$15 in cloud compute and 20 hours of work. Final editing and syncing were done in DaVinci Resolve.
More from Multimodal
- EG-FM Lowers FID to 1.45 in Pixel-Space Image Generation Without Backbone Changes — burny_tech · 2026-08-10
- Alibaba Releases Wan-Animate-2: Open-Source SOTA for Character Animation Transfer — natesiggard · 2026-08-10
- "Jim's Revenge": A Video Showcase Generated by Minimax H3 — blackdatafilms · 2026-08-10
- MiniMax-H3 Low-VRAM Setup: Save 10GB+ VRAM and Eliminate Swapping — Annual_Mess_1839 · 2026-08-10
- Musk Retweets Grok Imagine Test Showing Perfect Character Consistency in Video — elonmusk · 2026-08-10
- Spent $300 to Generate a Trailer for a REAL Game Using AI — KaLoiCus · 2026-08-10