Minimax H3 Tested: Subjects Speak Gibberish in Image-to-Video Generation

shadowmancer404 · reddit · 2026-08-05

A user utilizing a ComfyUI template reported an anomaly when using the Minimax H3 model to generate videos from a reference image: the subjects in the generated video uncontrollably make gibberish speaking noises.

The user noted that even when explicitly adding prompts like 'no talking' or 'no audio,' the model still fails to prevent the generation of these meaningless speaking actions and sounds. This exposes potential shortcomings in the current model's audio/motion control capabilities.

Related event: MiniMax H3 video generation glitches with unwanted dialogue(3 posts)→

Original post →

More from Multimodal

Multimodal channel →