Grok Imagine in Action: A Textbook Structured Prompt for Text-to-Video

tetsuoai · x · 2026-08-22

A user shared the full text-to-video prompt written for Grok Imagine: a 15-second locked-off shot in a cluttered gaming apartment where an anime woman knocks out a gamer with a frying pan and runs off giggling. The value lies in the structure: the prompt is broken into scene context, a location map (foreground/midground/background), first frame & blocking, format mode (one continuous shot, no cuts), optics (47° wide FOV, deep focus, no drift), and camera (locked tripod), followed by a timeline describing each action beat to 0.5-second precision (sneak-up → wind-up → 0.4s swing at 30 km/h → knockout). A solid reference template for text-to-video prompt engineering.

Original post →

More from Multimodal

Multimodal channel →