MiniMax H3 Ref2V Workflow: Face Swaps, Pose Transfer, Multi-Subject Mixes in One Chain
TBG______ · reddit · 2026-09-01
Building on an earlier lip-sync demo, the author built a general editing chain with MiniMax H3's Ref2V that treats a source clip as the "performance master" (motion, timing, camera) while identity/appearance is driven by reference images and audio.
Verified combinations:
- Music-video pose transfer
- Single character swap (main performer → reference character)
- Multi-subject mixes: identity swaps, 2-3 characters acting in/out of sync, pulling a character from the reference video into an image scene plus new characters
Key practical details:
- Everything runs through a single H3 chain with mixed references (ref-image + ref-video + ref-audio), using structured prompts that separate identity (image), performance (video), and constraints (text)
- A subject-definition template is provided, e.g. <Subject 1> must 100% match <Picture 1> in shape, material, logos and stay legible throughout
- Not cheap: reference-video conditioning pushed VRAM past 40GB on a 3-second, 2MP run; at least 2MP is needed for good detail and motion transfer
- The author notes this is just one more option among many and far from perfect
Related event: Community ComfyUI Workflows Unlock MiniMax H3 Video Editing Power(3 posts)→
More from Multimodal
- ECCV 2026 Tutorial: Post-Training Alignment and Enhancement for Diffusion Models — RisingSayak · 2026-09-01
- Clip Bot: Extract Captioned Highlights from YouTube Videos Automatically — CodeByPoonam · 2026-09-01
- Worldfold adds volumetric fog with light scattering, runs on M1 Mac — jamestagg · 2026-09-01
- Runway Omni 1.1 Flash hands-on: Enhanced control with first/last frames — jnack · 2026-09-01
- Scanning Puerto Rico: 1.39GB Gaussian Splats on WebGPU — willeastcott · 2026-09-01
- Creator shares 'Transformers plane' AI video, their best work to date — Such_Highlight_3941 · 2026-09-01