Reviewing AI video with separate checks for motion, subject, and background via VL models
professr_dumbledore · reddit · 2026-10-03
A concrete workflow for reviewing AI-generated video with a vision-language model, using Cosmos3-Edge for generation and Ling 3.0 VL reading six sampled frames against the intended motion.
- The example runs two generation instructions — an initial climb-then-right plan, then a Take B with hover 0–5s and drift right 5–10s — with cyan guides at fixed coordinates comparing start and end frames.
- Core method: split image-to-video review into three separate questions:
- Motion: where is the subject at start and end, on the same reference axes?
- Subject: which identifying details remain visible across samples?
- Background: how do background landmarks line up frame to frame?
- Separate questions make the next prompt revision specific: direction maps to the path, subject to appearance, background to the scene.
More from Multimodal
- PDMD 4-NFE LoRA for MiniMax-H3 released on Hugging Face, distills video diffusion to 4 steps — linoy_tsaban · 2026-10-04
- 2-step PDMD H3 LoRA tested: impressive for 2 steps, but H3 Turbo wins close-ups and speech — linoy_tsaban · 2026-10-04
- AI-generated Japanese town spawns unprompted interiors, including an art studio no one asked for — Merzmensch · 2026-10-04
- Handing off AI-generated images as decomposed RGBA layers for designer edits — tinkerbellyie · 2026-10-04
- Fan uses AI to put Asmongold into World of Warcraft and plays him — djcows · 2026-10-04
- POV: You're the Monkey King — AI video clip goes viral on Reddit — lucky-plume · 2026-10-04