MiniMax-H3 video inpainting ported to diffusers modular blocks: 6-step subject swap

linoy_tsaban · x · 2026-08-22

Linoy Tsaban ported MiniMax-H3's masked video and audio inpainting into Diffusers modular custom blocks (modular-diffusers), adapted from community ComfyUI workflows and released on Hugging Face under Apache-2.0.

In testing (with a turbo LoRA), a single reference photo plus 6 inference steps swaps the animal in a clip for the reference subject while the forest, snow, camera push, and original soundtrack stay untouched. The post includes a full Python example—loading blocks, passing mask/sourcevideo/sourceaudio and an image reference—plus a variant that skips the text encoder for split deployments.

Original post →

More from coding & agent

coding & agent channel →