Ref2VA face swap keeps the original face: four identity photos lose to the driving video

Motion16AI · reddit · 2026-09-20

Reddit user Motion16AI reports a face-swap failure with MiniMax H3's Ref2VA: rendering a 10-second video with four target-identity photos, the output keeps the original woman almost unchanged.

Setup: minimaxh3ref2vaprunedint8convrot, Heretic INT8 text encoder, 736×736, 240 frames at 24fps, 8-step MiniMaxH3TurboSampler, max-size reference images, video shift 12 / audio shift 3, memory-sparse optimization, no SAM/mask, face-swap adapter, or character LoRA.

The author suspects the full reference video is being treated as appearance conditioning rather than motion-only, overpowering the identity images, and is weighing pose extraction, fewer video reference frames, a corrected first frame, or a separate tracked face-swap stage. Asking the community for help.

Original post →

More from Multimodal

Multimodal channel →