A detailed ChatGPT-to-Kling-to-Suno workflow turns a Mumbai painting into a 15-second video
CurieuxExplorer · x · 2026-07-23
A creator shared a multi-tool image-to-video workflow for turning a rainy Mumbai painting into a cinematic clip.
The stack is:
- ChatGPT for prompt shaping
- Kling AI [I2V] for image-to-video generation
- Suno for music
- CapCut for editing
The prompt is unusually detailed: it asks the model to preserve the original composition exactly, keep the architecture and street geometry intact, and animate only the storm clouds and the heritage monument with subtle Van Gogh-like motion. The goal is a 15-second, 9:16 monsoon sequence that feels mostly photographic, with controlled painterly motion layered on top.
More from Multimodal
- Alibaba launches Qwen-Audio-3.0-TTS with 16 languages and 3-minute one-pass audio — Alibaba_Qwen · 2026-07-23
- Kling AI is said to handle close-up facial expressions better — burny_tech · 2026-07-23
- A reusable ChatGPT image prompt for a realistic portrait plus doodle-shadow twin — SimplyAnnisa · 2026-07-23
- Qwen-Image-3.0 aims for practical image generation, but still needs prompt tuning — 量子位 · 2026-07-23
- User showcases LTX 2.3 animations with a cinematic ogre-at-the-cake scene — Wise_Revolution385 · 2026-07-23
- MineExplorer shows top multimodal models collapse on long-horizon open-world tasks — 美团技术团队 · 2026-07-23