Minimax H3 Tested: Subjects Speak Gibberish in Image-to-Video Generation
shadowmancer404 · reddit · 2026-08-05
A user utilizing a ComfyUI template reported an anomaly when using the Minimax H3 model to generate videos from a reference image: the subjects in the generated video uncontrollably make gibberish speaking noises.
The user noted that even when explicitly adding prompts like 'no talking' or 'no audio,' the model still fails to prevent the generation of these meaningless speaking actions and sounds. This exposes potential shortcomings in the current model's audio/motion control capabilities.
Related event: MiniMax H3 video generation glitches with unwanted dialogue(3 posts)→
More from Multimodal
- xllm generates an image in 0.4 seconds — warycat · 2026-08-24
- Wan 3.0 Video Generation Demo: Single Prompt, First Attempt — OdinLovis · 2026-08-24
- Wan 3.0 launches on fal with native 30-second video generation — aziz4ai · 2026-08-24
- Animation workflow combining LTX2.3 FLFA2V and Suno Audio — JustRuss79 · 2026-08-24
- Minimax H3 T2VA supports 15+ characters simultaneously on screen — beatlepol · 2026-08-24
- How to replace an image element with Flux Kontext or Qwen Image Edit? — Wemos_D1 · 2026-08-24