MiniMax video generation: how prompt behavior drifts with resolution, length and sampler

boriskarloff83 · reddit · 2026-10-09

The author is building "plug-and-play" batch video generation with MiniMax: swap any reference image (red VW, blue BMW) into the same baseline prompt for a car-crash scenario, and get consistent behavior. But baseline prompts prove hard to stabilize: prompt adherence differs between 0.4MP and 0.5MP, adding one second of length changes behavior completely, samplers/schedulers alter prompt semantics beyond visual fidelity, LoRAs bring side effects, and prompt timestamps interact with everything — while image models were far more controllable. The post also covers pose control from references and seeks others' iteration workflows.

Original post →

More from Multimodal

Multimodal channel →