Claude, Suno and Seedance turned one idea into 38 clips for about $87
MosskeepForest · reddit · 2026-07-25
The author describes a fast, mostly hands-off multimodal video workflow.
Pipeline
- Claude was used to write the song from a rough premise
- Suno was used to try out the song quickly
- A directing skill then converted the lyrics and song length into 15-second chunks, shots, environments, and character lists
- GPT was used to generate the asset list
- Seedance 2 produced the clips
Results
- 38 clips in total
- about 2 to 3 hours end-to-end
- roughly $87 USD in production cost
The author notes that they did not spend much time reading back the shot plan, specifically to test how fast the workflow could run.
More from Multimodal
- Apple-PI benchmarks video-model reasoning against explicit physical laws — _akhaliq · 2026-07-25
- LTX2.3 can render a 20-second video with audio on a 7800XT in 12.5 minutes — okfine1337 · 2026-07-25
- Inflect-Micro-v2 trends on Hugging Face as a small local TTS model — owensong · 2026-07-25
- PixelRAG skips HTML parsing, uses screenshots for web retrieval, and beats text RAG by 18.1% — Roger_M_Taylor · 2026-07-25
- Skywork Video pushes storyboard-first AI video production as its core upgrade — Shruti_0810 · 2026-07-25
- Wan SCAIL-2 adds segmentation control with better auto-extend and less color shift — External_Trainer_213 · 2026-07-25