One-Person AI Music Video Took a Month: 20 Stills and 10 Takes Per Clip
kraussian · reddit · 2026-09-23
A creator spent a month producing a full AI-generated music video using Suno for music, GPT Image for stills, and MiniMax H3 (checkpoint + Turbo LoRA) in ComfyUI for animation, with Claude helping write a stitching script. Each 5-10s clip averaged 20 still iterations and 10 animated takes. He calls the result 90% there, notes classic MiniMax tells (blurred distant faces, hallucinated motion), and says the 'one-person film studio' is less than a year away.
More from Multimodal
- Testing Hunyuan Image 3.5 via Miora's Image Generator workflow — HeyAmit_ · 2026-09-23
- Hunyuan Image 3.5 realism test: lighting, materials and camera language — HeyAmit_ · 2026-09-23
- Hands-on: Tencent's Hunyuan Image 3.5 preview impresses with text, references and editing control — HeyAmit_ · 2026-09-23
- NetEase Youdao open-sources 2B streaming ASR and 14B Chinese-English simultaneous translation models — AdinaYakup · 2026-09-23
- Opus 5.5 one-shots a full newsletter launch video, creator says "it's over for video guys" — alex_verem · 2026-09-23
- PixVerse R2 hands-on: AI video that becomes a world you walk through with WASD — HeyAmit_ · 2026-09-23