MiniMax H3 workflow: Image-to-video with integrated prompt enhancer

Friendly-Fig-6015 · reddit · 2026-09-01

A user shared a MiniMax H3 Image-to-Video workflow with an integrated prompt enhancer. It uses a VLM to analyze reference images and simple descriptions, generating detailed prompts optimized for H3 (characters, actions, camera, sound). Tests show significantly better instruction adherence, especially for complex actions and dialogue.

Original post →

More from Multimodal

Multimodal channel →