Creator renders AI music video with local video model after 48 hours on two GPUs
tetsuoai · x · 2026-10-10
Creator tetsuoai shared an almost fully AI-produced music video built on Scott Buckley's "Titan": he transcribed the melody by ear into a piano roll, then layered pads, vocals, synths, and bass using Suno.
Every video frame came from a local video model that ran his two GPUs at full load for over 48 hours straight; he notes he was out of heavy usage and that Grok Imagine would have been the easier route. Overlays and post-processing were handled by Grok and Opus.
Even the "inside of a grand piano" shot is synthetic: a render that reads MIDI and path-traces everything in Blender, cooked with Grok + Opus 5.5. A concrete end-to-end pipeline combining a local video model with multi-model collaboration.
More from Multimodal
- Generating a themed Sakura pagoda garden LEGO image with fal — noahsolomon · 2026-10-10
- LLMs writing code beat Seedance at some video types, blogger observes — gefei55 · 2026-10-10
- "Game over": X user shares what he calls the most realistic AI video yet — alexgoughcooper · 2026-10-10
- Unsloth releases Qwen-Image-2.1-Turbo-FP8, quantized text-to-image and image editing — unsloth · 2026-10-10
- Qwen Image 2.1 vs Turbo: a fairer comparison with proper sampler settings — No-Zookeepergame4774 · 2026-10-10
- Midjourney starts testing its MCP with a limited creative community — midjourney · 2026-10-10