Swap any video subject with one image: Viggle-Animate (MiniMax H3 finetune) needs no pose detection, 3 inference steps

cocktailpeanut · x · 2026-09-15

A developer demos Viggle-Animate, a MiniMax H3 finetune for video subject swapping: extract the first frame, use GPT Image 2.5 to swap characters, then feed that single reference image plus the original video into the model — it syncs the reference to the video till the end. No pose detection, no SAM, no segmentation pipeline; multiple subjects can be swapped at once, and it runs in just 3 inference steps. The author calls it criminally underrated.

Related event: Maestro Local AI Studio and Video Character Swap Workflows Gain Attention(3 posts)→

Original post →

More from Multimodal

Multimodal channel →