Fixing AI video face melt: upload raw audio and quote dialogue in prompt
Fragrant-Cheek-4273 · reddit · 2026-08-31
The author encountered face melting in AI videos using Minimax H3, attributing it to independent VAEs for audio and visual data. Audio data is less dense, causing the audio encoder to overfit and pull spatial anchors out of alignment.
Fix: upload the raw audio reference file and type the exact dialogue in double quotation marks inside the positive text prompt. This forces the model to map lip movements to the literal text string while referencing the audio VAE, stabilizing facial expressions. Not perfect, but passable.
More from Multimodal
- LichtFeld Studio: Open-Source 3D Gaussian Splatting with Training, Editing, and MCP Automation — janusch_patas · 2026-08-31
- Seedance 2.5 Prompt Puts Meme Characters in One 30s Unbroken NYC Vlog Take — socialwithaayan · 2026-08-31
- Every Meme You Know Comes Alive in One Unbroken NYC Vlog via Seedance 2.5 — socialwithaayan · 2026-08-31
- Why Flux slows down on an 8GB card: ComfyUI silently spills to CPU when VRAM fills — leonbuilds · 2026-08-31
- Seedance 2.5 Generates Photorealistic IGI-Style Tactical FPS Gameplay — SimplyAnnisa · 2026-08-31
- Lovart launches China edition: one agent delivers full brand design kits, reference plugin and 100 pro skills — 卡尔的AI沃茨 · 2026-08-31