Using Qwen3-VL-4B to auto-expand prompts for Minimax H3 video generation

CountFloyd_ · reddit · 2026-09-04

The author chains a Generate Text node with the qwen3vl4b CLIP model, feeding it Minimax's prompting guide to auto-expand bland base prompts into rich multi-shot prompts for H3 video generation — with better-than-expected results. Remaining issue: the small model consistently ignores the 5-second per-shot duration cap even when it's in the system prompt.

Original post →

More from Multimodal

Multimodal channel →