A practical first-frame anchor prompt template for MiniMax H3 image-to-video

VasaFromParadise · reddit · 2026-08-17

A Reddit user shares a prompt structure that stabilizes MiniMax H3 image-to-video generations: organize descriptions as first-frame anchor → action onset → continuous development → result/reaction, first locking character identity, clothing, colors, composition and lighting anchors before describing motion, keeping consistency across the clip.

The workflow uses an LLM to auto-generate the prompt from the input image, produces a 0.5-megapixel 7-second video, then upscales with RTX Video Super Resolution to 1.5x and interpolates to 48fps — runnable on any PC, each step taking under a minute. The example walks through a red-haired woman turning toward the camera with a surprised smile and speaking.

Original post →

More from Multimodal

Multimodal channel →