Pablo Stanley shares a full AI video workflow using ChatGPT, Gemini, Runway and CapCut
jdjohnson · x · 2026-07-21
Pablo Stanley breaks down his AI video workflow:
- He starts by generating still frames in ChatGPT, which he says works best for preserving the same real person across different angles and styles.
- He then feeds those frames into Gemini’s video tool, writing a script first and iterating on directing prompts because the default voice tone is too chirpy.
- For his own shots, he records live footage first and uses Runway’s Aleph 2 for video-to-video conversion, with the ChatGPT image as a style reference.
- Finally, he assembles everything in CapCut.
The post also includes a snapshot showing the same man rendered in two different video/image setups, highlighting how the workflow keeps character identity consistent.
More from Multimodal
- TimeLens2 claims SOTA on 7 video grounding benchmarks with 4B and 8B models — _akhaliq · 2026-07-21
- AI-made 4-minute horror short ‘THE NOT KNOW’ lands as a shareable demo — gen_ericai · 2026-07-21
- SVG Generation Comparison: Leading AI Models Draw a Red Ferrari — Able-Line2683 · 2026-07-21
- Adding order metadata makes VLM error detection collapse, new benchmark shows — m_wulfmeier · 2026-07-21
- Claude AGI Agent starts paging itself in Slack with a no-heartbeat alarm — Sauers_ · 2026-07-21
- Gemini Omni is being called a video-editing leap on par with Nano Banana — CodeByPoonam · 2026-07-21