Fixing AI video face melt: upload raw audio and quote dialogue in prompt

Fragrant-Cheek-4273 · reddit · 2026-08-31

The author encountered face melting in AI videos using Minimax H3, attributing it to independent VAEs for audio and visual data. Audio data is less dense, causing the audio encoder to overfit and pull spatial anchors out of alignment.

Fix: upload the raw audio reference file and type the exact dialogue in double quotation marks inside the positive text prompt. This forces the model to map lip movements to the literal text string while referencing the audio VAE, stabilizing facial expressions. Not perfect, but passable.

Original post →

More from Multimodal

Multimodal channel →