Minimax H3 users struggle to keep character dialogue turns consistent in generated videos

Efficient_Heron5978 · reddit · 2026-09-18

A user generating videos with Minimax H3 (R2V) on ComfyUI cloud reports that even with shot-by-shot prompts specifying which character speaks each line, dialogue frequently gets assigned to the wrong character, spoken by one character alone, or delivered by both simultaneously.

The post shows the full prompt structure — reference images, camera, background sounds, per-line tone and accent notes — and asks how to achieve consistent speaker turns in AI video generation.

Original post →

More from Multimodal

Multimodal channel →